VLDB 2026 Research / reviewers in the wild / expert
Hong Zhang 0013
dblp:24/6914-13
· DBLP profile ↗
202ranked-venue papers
18as first author
53since 2021 · last 2026
0000-0002-1677-6132ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 142 · 13 first-author · 34 since 2021Systems, architecture and hardware · 106 · 11 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing Effective and Efficient Category-Level Articulated Object Pose PerceptionabstractCategory-level articulated object pose perception-encompassing both static pose estimation and dynamic pose tracking-is critical for embodied AI systems interacting with complex environments. Due to the inherent complexity and diverse motion structures of articulated objects, existing methods often exhibit limitations in adequately modeling kinematic constraints, handling self-occlusions, and meeting optimization requirements. Building upon EfficientCAPER (Yu et al., 2024), this work introduces CAPER++, a unified framework addressing these limitations through three key innovations: first, a joint-centric hierarchical model decomposes objects into a root part and constrained parts linked by joints, explicitly embedding kinematic constraints for geometrically consistent pose recovery. Second, an SE(3) manifold formulation leverages Lie algebra in the tangent space for singularity-free rotation representation and stable optimization, replacing error-prone direct regression. Third, for tracking, a proxy canonicalization strategy reformulates pose updates as SE(3) increment predictions relative to keyframes, enhanced by a dynamic keyframe mechanism to suppress drift. Extensive experiments on synthetic (ArtImage, PM-Videos), semi-synthetic (ReArtMix, ReArt-Videos), and real-world (RobotArm, RobotArm-Videos) benchmarks demonstrate state-of-the-art accuracy and robustness. CAPER++ achieves real-time inference (50 FPS) without post-processing, significantly advancing category-level articulated perception for real-world applications. Li Zhang 0104, Xianhui Meng, Liu Liu 0012, Rujing Wang, Cewu Lu, Jun Liu 0004, Hong Zhang 0013 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | Dexterous Manipulation Through Imitation Learning: A SurveyabstractDexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain. Shan An, Chao Tang 0001, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu 0001, Ran Song 0001, Wei Zhang 0021, Zeng-Guang Hou, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 12 |
| 2026 | Accurate and Robust UWB Localization With Incomplete Measurements Based on Multi-Modal Diffusion Model
Ming Sun 0025, Bo Yang 0019, Li He 0002, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | ROVER: Robust Loop Closure Verification With Trajectory Prior in Repetitive EnvironmentsabstractLoop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical keyframes, achieving drift correction and global relocalization. However, a falsely detected loop can be fatal, and this is especially difficult in repetitive environments where appearance-based features fail due to the high similarity. Therefore, verifying a loop closure is a critical step to avoid false-positive detections. Existing works in loop closure verification predominantly focus on learning invariant appearance features, neglecting the prior knowledge of the robot’s spatial-temporal motion cue, i.e., trajectory. In this article, we propose ROVER, a loop closure verification method that leverages the historical trajectory as a prior constraint to reject false loops in challenging repetitive environments. For each loop candidate, it is first used to estimate the robot trajectory with pose-graph optimization. This trajectory is then submitted to a scoring scheme that assesses its compliance with the trajectory without the loop, which we refer to as the trajectory prior constraint (TPC), to determine if the loop candidate should be accepted. Benchmark comparisons and real-world experiments demonstrate the effectiveness of the proposed method. Furthermore, we integrate ROVER into state-of-the-art SLAM systems to verify its robustness and efficiency. Jingwen Yu, Jianhao Jiao, Anjun Hu, Zhonghang Liu, Jiankun Wang 0001, Ping Tan 0002, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2026 | Two-Step Nyström Sampling for Large-Scale Kernel ApproximationabstractNystrom approximation is one of the most popular approximation methods to accelerate kernel analysis on largescale data sets. Nystrom employs one single landmark set to ¨ obtain eigenvectors (low-rank decomposition) and projects the entire data set to the eigenvectors (embedding). Most existing methods focus on accelerating landmark selection. For extremely large-scale data sets, however, the embedding time cost, rather than that of low-rank decomposition, is critical. In addition, both accuracy and embedding time cost are dominated by the landmark set size. As a result, using more landmarks is the only way to improve accuracy at the cost of extremely high embedding costs. In this paper, we propose a method for the first time to decouple embedding cost from that of low-rank decomposition. We first obtain the eigenvectors from a large landmark set for a low error, and then optimize a small landmark set that minimizes the landmark-set-embedding error to ensure a low embedding cost. In return, our accuracy is close to that of the large landmark set but the small one dominates the embedding time cost. Our method can deal with popular kernels and be plugged into most existing methods. Experimental results demonstrate the superiority of the proposed method. Li He 0002, Hong Zhang 0013 |
IEEE Trans. Big Data | 2 |
| 2025 | RTAGrasp: Learning Task-Oriented Grasping from Human Videos via Retrieval, Transfer, and AlignmentabstractTask-oriented grasping (TOG) is crucial for robots to accomplish manipulation tasks, requiring the determination of TOG positions and directions. Existing methods either rely on costly manual TOG annotations or only extract coarse grasping positions or regions from human demonstrations, limiting their practicality in real-world applications. To address these limitations, we introduce RTAGrasp, a Retrieval, Transfer, and Alignment framework inspired by human grasping strategies. Specifically, our approach first effortlessly constructs a robot memory from human grasping demonstration videos, extracting both TOG position and direction constraints. Then, given a task instruction and a visual observation of the target object, RTAGrasp retrieves the most similar human grasping experience from its memory and leverages semantic matching capabilities of vision foundation models to transfer the TOG constraints to the target object in a training-free manner. Finally, RTAGrasp aligns the transferred TOG constraints with the robot's action for execution. Evaluations on the public TOG benchmark, TaskGrasp dataset, show the competitive performance of RTAGrasp on both seen and unseen object categories compared to existing baseline methods. Real-world experiments further validate its effectiveness on a robotic arm. Our code, appendix, and video are available at https://sites.google.com/view/rtagrasp/home. Wenlong Dong, Dehao Huang, Jiangshan Liu, Chao Tang 0001, Hong Zhang 0013 |
ICRA | 5 |
| 2025 | FLAF: Focal Line and Feature-Constrained Active View Planning for Visual Teach and RepeatabstractThis paper presents FLAF, a focal line and feature-constrained active view planning method for autonomous orientation adjustment of a rotatable active camera during mobile robot navigation. FLAF is built on a visual teach-and-repeat (VT&R) system, which enables robots to cruise various paths that fulfill many daily autonomous navigation requirements. The VT&R system integrates Visual Simultaneous Localization and Mapping (VSLAM) with trajectory following. However, tracking failures in feature-based VSLAM, particularly in textureless regions common in human-made environments, poses a significant challenge to real-world VT&R deployment. To address this, the proposed view planner is integrated into a feature-based VSLAM system, creating an active camerabased VSLAM (AC-SLAM) solution that mitigates tracking failures. Our system features a Pan-Tilt Unit (PTU)-based active camera mounted on a mobile robot. FLAF actively directs the camera toward more map points during path learning and toward more feature-identifiable map points while following the learned trajectory. Using FLAF, the AC-SLAM system constructs a complete path map during teaching and maintains stable localization during repeating. Experimental results in real scenarios show that FLAF significantly outperforms existing methods by accounting for feature identifiability, particularly the view angle of the features. During effectively dealing with low-texture regions in active view planning, considering feature identifiability enables our active VT&R system to perform well in challenging environments. Changfei Fu, Weinan Chen, Wenjun Xu 0005, Hong Zhang 0013 |
ICRA | 4 |
| 2025 | Optimizing NeRF-Based SLAM with Trajectory Smoothness ConstraintsabstractThe joint optimization of Neural Radiance Fields (NeRF) and camera trajectories has been widely applied in SLAM tasks due to its superior dense mapping quality and consistency. NeRF-based SLAM learns camera poses using constraints by implicit map representation. A widely observed phenomenon that results from the constraints of this form is jerky and physically unrealistic estimated camera motion, which in turn affects the map quality. To address this deficiency of current NeRF-based SLAM, we propose in this paper TS-SLAM (TS for Trajectory Smoothness). It introduces smoothness constraints on camera trajectories by representing them with uniform cubic B-splines with continuous acceleration that guarantees smooth camera motion. Benefiting from the differentiability and local control properties of B-splines, TS-SLAM can incrementally learn the control points end-to-end using a sliding window paradigm. Additionally, we regularize camera trajectories by exploiting the dynamics prior to further smooth trajectories. Experimental results demonstrate that TS-SLAM achieves superior trajectory accuracy and improves mapping quality versus NeRF-based SLAM that does not employ the above smoothness constraints. Yicheng He, Guangcheng Chen, Hong Zhang 0013 |
ICRA | 3 |
| 2025 | HGDiffuser: Efficient Task-Oriented Grasp Generation via Human-Guided Grasp Diffusion ModelsabstractTask-oriented grasping (TOG) is essential for robots to perform manipulation tasks, requiring grasps that are both stable and compliant with task-specific constraints. Humans naturally grasp objects in a task-oriented manner to facilitate subsequent manipulation tasks. By leveraging human grasp demonstrations, current methods can generate high-quality robotic parallel-jaw task-oriented grasps for diverse objects and tasks. However, they still encounter challenges in maintaining grasp stability and sampling efficiency. These methods typically rely on a two-stage process: first performing exhaustive task-agnostic grasp sampling in the 6-DoF space, then applying demonstration-induced constraints (e.g., contact regions and wrist orientations) to filter candidates. This leads to inefficiency and potential failure due to the vast sampling space. To address this, we propose the Human-guided Grasp Diffuser (HGDiffuser), a diffusion-based framework that integrates these constraints into a guided sampling process. Through this approach, HGDiffuser directly generates 6-DoF task-oriented grasps in a single stage, eliminating exhaustive task-agnostic sampling. Furthermore, by incorporating Diffusion Transformer (DiT) blocks as the feature backbone, HGDiffuser improves grasp generation quality compared to MLP-based methods. Experimental results demonstrate that our approach significantly improves the efficiency of task-oriented grasp generation, enabling more effective transfer of human grasping strategies to robotic systems. To access the source code and supplementary videos, visit https://sites.google.com/ view/hgdiffuser. Dehao Huang, Wenlong Dong, Chao Tang 0001, Hong Zhang 0013 |
IROS | 4 |
| 2025 | Reducing Redundancy in VSLAM: VLMs-driven Keyframe Selection using Multi-dimensional Semantic InformationabstractKeyframe selection plays a crucial role in balancing computational efficiency and localization accuracy in Visual Simultaneous Localization and Mapping (VSLAM) systems. Existing keyframe selection methods often struggle to capture high-level semantic information in environments where multiple semantic dimensions interact. In this paper, we propose the Multi-dimensional Semantic Analysis (MSA) module based on Visual-Language Models (VLMs). By leveraging the capability of VLMs to extract rich semantic features, we compute the similarity between each image frame and a set of textual descriptions, generating a scene descriptor that quantifies the semantic distance between frames across multiple dimensions (e.g., object count, texture, and lighting). We then introduce the Scene Change Assessment (SCA) module based on Bayesian On-line Changepoint Detection (BOCD), which identifies keyframes with significant semantic information gain, thereby reducing the total number of keyframes. Extensive experiments on an open dataset demonstrate that our method not only significantly reduces the number of keyframes but also maintains high localization accuracy. Furthermore, the inference speed of the MSA module satisfies the real-time requirements of VSLAM. These results underscore the potential of our approach to enhance the efficiency of keyframe selection. Xiang Huo, Shilang Chen, Haifei Zhu, Yisheng Guan, Hong Zhang 0013, Weinan Chen |
IROS | 6 |
| 2025 | JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent InteractionabstractPredicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability modes in multi-agent joint prediction. To tackle this issue, we propose a two-stage multi-agent interactive prediction framework named keypoint-guided joint prediction after classification-aware marginal proposal (JAM). The first stage is modeled as a marginal prediction process, which classifies queries by trajectory type to encourage the model to learn all categories of trajectories, providing comprehensive mode information for the joint prediction module. The second stage is modeled as a joint prediction process, which takes the scene context and the marginal proposals from the first stage as inputs to learn the final joint distribution. We explicitly introduce key waypoints to guide the joint prediction module in better capturing and leveraging the critical information from the initial predicted trajectories. We conduct extensive experiments on the real-world Waymo Open Motion Dataset interactive prediction benchmark. The results show that our approach achieves competitive performance. In particular, in the framework comparison experiments, the proposed JAM outperforms other prediction frameworks and achieves state-of-the-art performance in interactive trajectory prediction. The code is available at https://github.com/LinFunster/JAM to facilitate future research. Fangze Lin, Ying He 0006, F. Richard Yu, Hong Zhang 0013 |
IROS | 4 |
| 2025 | FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction FollowingabstractRobotic instruction following tasks require seamless integration of visual perception, task planning, target localization, and motion execution. However, existing task planning methods for instruction following are either data-driven or underperform in zero-shot scenarios due to difficulties in grounding lengthy instructions into actionable plans under operational constraints. To address this, we propose FlowPlan, a structured multi-stage LLM workflow that elevates zero-shot pipeline and bridges the performance gap between zero-shot and data-driven in-context learning methods. By decomposing the planning process into modular stages—task information retrieval, language-level reasoning, symbolic-level planning, and logical evaluation—FlowPlan generates logically coherent action sequences while adhering to operational constraints and further extracts contextual guidance for precise instance-level target localization. Benchmarked on ALFRED and validated in real-world applications, our method achieves competitive performance relative to data-driven in-context learning methods and demonstrates adaptability across diverse environments. This work advances zero-shot task planning in robotic systems without reliance on labeled data. Project website: https://instruction-following-project.github.io/. Chao Tang 0001, Hanjing Ye, Hong Zhang 0013 |
IROS | 4 |
| 2025 | TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and VerificationabstractVisual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the highly repetitive structures inherent in such environments. We observe that scene texts frequently appear in indoor spaces and can help distinguish visually similar but different places. This inspires us to propose TextInPlace, a simple yet effective VPR framework that integrates Scene Text Spotting (STS) to mitigate visual perceptual ambiguity in repetitive indoor environments. Specifically, TextInPlace adopts a dual-branch architecture within a local parameter sharing network. The VPR branch employs attention-based aggregation to extract global descriptors for coarse-grained retrieval, while the STS branch utilizes a bridging text spotter to detect and recognize scene texts. Finally, the discriminative texts are filtered to compute text similarity and re-rank the top-K retrieved images. To bridge the gap between current text-based repetitive indoor scene datasets and the typical scenarios encountered in robot navigation, we establish an indoor VPR benchmark dataset, called Maze-with-Text. Extensive experiments on both custom and public datasets demonstrate that TextInPlace achieves superior performance over existing methods that rely solely on appearance information. The dataset, code, and trained models are publicly available at https://github.com/HqiTao/TextInPlace. Huaqi Tao, Bingxi Liu 0001, Calvin Chen, Tingjun Huang, Jinqiang Cui, Hong Zhang 0013 |
IROS | 7 |
| 2025 | Heterogeneous Graph Network-Based UWB Localization for Complex Indoor EnvironmentsabstractAccurate indoor location-based services are important for mobile robots, especially in complex indoor environments. In this paper, we propose a heterogeneous graph network-based ultra-wide band (UWB) localization method to provide accurate and robust localization results for mobile robots in complex indoor scenarios. The core of our approach lies in constructing the anchors, ranging measurements and tags into a heterogeneous graph structure according to the topological structure of the UWB localization system, and then design a spatial-temporal heterogeneous graph attention neural network to extract high-level features and estimate the tag locations from the graph. Therefore, the geometric relationships contained in the UWB localization system are comprehensively established, while the spatial and temporal information contained in the ranging measurements can also be extracted. We validate the proposed method through real-world experiments. The results demonstrate that, compared to existing deep learning-based methods, the constructed heterogeneous graph better represents the geometric structure of the UWB localization system, and the designed heterogeneous graph neural network effectively extracts the spatial-temporal and geometric features. Consequently, the accuracy and robustness of UWB localization are significantly improved. Bo Yang 0019, Sizhen He, Weinan Chen, Hong Zhang 0013 |
IROS | 5 |
| 2025 | Monocular Person Localization under Camera Ego-MotionabstractLocalizing a person from a moving monocular camera is critical for Human-Robot Interaction (HRI). To estimate the 3D human position from a 2D image, existing methods either depend on the geometric assumption of a fixed camera or use a position regression model trained on datasets containing little camera ego-motion. These methods are vulnerable to fierce camera ego-motion, resulting in inaccurate person localization. We consider person localization as a part of a pose estimation problem. By representing a human with a four-point model, our method jointly estimates the 2D camera attitude and the person’s 3D location through optimization. Evaluations on both public datasets and real robot experiments demonstrate our method outperforms baselines in person localization accuracy. Our method is further implemented into a person-following system and deployed on an agile quadruped robot. Hanjing Ye, Hong Zhang 0013 |
IROS | 3 |
| 2025 | X-RepSLAM: VLM-Driven Adaptive Cross-Representation Visual SLAMabstractVisual Simultaneous Localization and Mapping (VSLAM) is a critical technology for autonomous driving and mobile robotics. Traditional VSLAM methods based on discrete representations, such as point clouds, offer high computational efficiency and excellent localization accuracy, but they exhibit limited robustness. In contrast, methods employing field representations, like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS), provide greater robustness at the expense of increased computational demands and reduced localization accuracy. Hybrid VSLAM approaches that attempt to combine these representations typically rely on serial, synchronous cascades, which compromise robustness, computational efficiency, and GPU memory usage. This paper introduces a novel adaptive cross-representation VSLAM framework that applies different representation modeling techniques to distinct regions of an image sequence and adopts asynchronous parallel modeling in overlapping regions. A Vision Language Model (VLM) is used to analyze the image sequence, enabling the detection of representation modeling regions and adaptive switching between representations. Cross-representation data association is performed through a coarse-to-fine feature selection process, resulting in a globally consistent map. The proposed method is evaluated on both public and custom-collected datasets, where experimental results show that it surpasses state-of-the-art methods in terms of robustness, computational efficiency, localization accuracy, and GPU memory usage. Shilang Chen, Sehua Ji, Xuefeng Zhou, Hong Zhang 0013, Weinan Chen, Yisheng Guan |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Edge-Assisted Multi-Robot Visual-Inertial SLAM With Efficient CommunicationabstractThe integration of cloud computing and edge computing is an effective way to achieve global consistent and real-time multi-robot Simultaneous Localization and Mapping (SLAM). Cloud computing effectively solves the problem of limited computing, communication and storage capacity of terminal equipment. However, limited bandwidth and extremely long communication links between terminal devices and the cloud result in serious performance degradation of multi-robot SLAM systems. To reduce the computational cost of feature tracking and improve the real-time performance of the robot, a lightweight SLAM method of optical flow tracking based on pyramid IMU prediction is proposed. On this basis, a centralized multi-robot SLAM system based on a robot-edge-cloud layered architecture is proposed to realize real-time collaborative SLAM. It avoids the problems of limited on-board computing resources and low execution efficiency of single robot. In this framework, only the feature points and keyframe descriptors are transmitted and lossless encoding and compression are carried out to realize real-time remote information transmission with limited bandwidth resources. This design reduces the actual bandwidth occupied in the process of data transmission, and does not cause the loss of SLAM accuracy caused by data compression. Through experimental verification on the EuRoC dataset, compared with the current most advanced local feature compression method, our method can achieve lower data volume feature transmission, and compared with the current advanced centralized multi-robot SLAM scheme, it can achieve the same or better positioning accuracy under low computational load.Note to Practitioners—The purpose of this paper is to reduce the communication load of a Cloud-Edge-Robot system by compressing and transmitting of keyframes and non-keyframes, respectively, which is suitable for a multi-robot SLAM system and can realize multi-robot joint localization and sparse map reconstruction under efficient communication. Currently, remote SLAM or centralized multi-robot SLAM is usually implemented by transferring the whole image or the features and descriptors of the image. In this paper, lightweight SLAM optical flow tracking based on pyramid IMU prediction is implemented to track non-keyframes. At the edge server, tracking between non-keyframes is realized only by transmitting keypoints. For keyframes, the pose estimation is realized by transmitting compressed features and descriptors. Multi-robot localization and map fusion are realized in the cloud through key frame feature information. Experiments on public datasets show that this method is feasible and can achieve high-precision joint positioning with a low amount of transmitted data. In future studies, we will apply this framework to more real-world systems, while achieving rich, accurate map fusion with more advanced features. Xin Liu 0068, Shuhuan Wen, Jing Zhao 0020, Tony Z. Qiu, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | FoundationGrasp: Generalizable Task-Oriented Grasping With Foundation ModelsabstractTask-oriented grasping (TOG), which refers to synthesizing grasps on an object that are configurationally compatible with the downstream manipulation task, is the first milestone towards tool manipulation. Analogous to the activation of two brain regions responsible for semantic and geometric reasoning during cognitive processes, modeling the intricate relationship between objects, tasks, and grasps necessitates rich semantic and geometric prior knowledge about these elements. Existing methods typically restrict the prior knowledge to a closed-set scope, limiting their generalization to novel objects and tasks out of the training set. To address such a limitation, we propose FoundationGrasp, a foundation model-based TOG framework that leverages the open-ended knowledge from foundation models to learn generalizable TOG skills. Extensive experiments are conducted on the contributed Language and Vision Augmented TaskGrasp (LaViA-TaskGrasp) dataset, demonstrating the superiority of FoundationGrasp over existing methods when generalizing to novel object instances, object classes, and tasks out of the training set. Furthermore, the effectiveness of FoundationGrasp is validated in real-robot grasping and manipulation experiments on a 7-DoF robotic arm. Our code, data, appendix, and video are publicly available athttps://sites.google.com/view/foundationgrasp. Note to Practitioners—This research is motivated by the challenge of generalizable task-oriented grasping skill learning. Solving such a challenge could significantly improve the robot’s level of automation and intelligence in tool manipulation for household and industrial tasks. Existing methods struggle with handling unseen objects and tasks in dynamic, open-world environments. To overcome this limitation, we propose to leverage the open-ended knowledge from foundation models to improve the generalization capabilities of existing TOG methods. This way, the robot can perform TOG w.r.t. unseen objects and tasks, facilitating downstream tool manipulation. Overall, this research has broad applicability to various scenarios involving tool manipulation, such as cleaning kitchenware and assembling parts in industrial contexts. Chao Tang 0001, Dehao Huang, Wenlong Dong, Ruinian Xu, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects
Guangcheng Chen, Yicheng He, Li He 0002, Hong Zhang 0013 |
ECCV (8) | 4 |
| 2024 | Efficient Object Rearrangement via Multi-view FusionabstractThe prospect of assistive robots aiding in object organization has always been compelling. In an image-goal setting, the robot rearranges the current scene to match the single image captured from the goal scene. The key to an image-goal rearrangement system is estimating the desired placement pose of each object based on the single goal image and observations from the current scene. In order to establish sufficient associations for accurate estimation, the system should observe an object from a viewpoint similar to that in the goal image. Existing image-goal rearrangement systems, due to their reliance on a fixed viewpoint for perception, often require redundant manipulations to randomly adjust an object’s pose for a better perspective. Addressing this inefficiency, we introduce a novel object rearrangement system that employs multi-view fusion. By observing the current scene from multiple viewpoints before manipulating objects, our approach can estimate a more accurate pose without redundant manipulation times. A standard visual localization pipeline at the object level is developed to capitalize on the advantages of multi-view observations. Simulation results demonstrate that the efficiency of our system outperforms existing single-view systems. The effectiveness of our system is further validated in a physical experiment. For videos, please visit https: //sites.google.com/view/multi-view-rearr. Dehao Huang, Chao Tang 0001, Hong Zhang 0013 |
ICRA | 3 |
| 2024 | A Point-to-distribution Degeneracy Detection Factor for LiDAR SLAM using Local Geometric ModelsabstractLimited by the working principles, LiDAR-SLAM systems suffer from the degeneration phenomenon in environments such as long corridors and tunnels, due to the lack of sufficient geometric features for frame-to-frame matching. The accuracy and sensitivity of existing degeneracy detection methods need to be further improved. In this paper, we propose a novel method for degeneracy detection using local geometric models based on point-to-distribution matching. To obtain an accurate description of local geometric models, an adaptive adjustment of voxel segmentation according to the point cloud distribution and density is designed. The codes of the proposed method is open-source and available at https://github.com/jisehua/Degenerate-Detection.git. Experiments with public datasets and self-build robots were conducted to evaluate the methods. The results exhibit that our proposed method achieves higher accuracy than the other existing approaches. Applying our proposed method is beneficial for improving the robustness of the LiDAR-SLAM systems. Sehua Ji, Weinan Chen, Zerong Su, Yisheng Guan, Jiehao Li, Hong Zhang 0013, Haifei Zhu |
ICRA | 6 |
| 2024 | Commonsense Scene Graph-based Target Localization for Object SearchabstractObject search is a fundamental skill for household robots, yet the core problem lies in the robot’s ability to locate the target object accurately. The dynamic nature of household environments, characterized by the arbitrary placement of daily objects by users, makes it challenging to perform target localization. To efficiently locate the target object, the robot needs to be equipped with knowledge at both the object and room level. However, existing approaches rely solely on one type of knowledge, leading to unsatisfactory object localization performance and, consequently, inefficient object search processes. To address this problem, we propose a commonsense scene graph-based target localization, CSG-TL, to enhance target object search in the household environment. Given the pre-built map with stationary items, the robot models the room-level spatial knowledge with object-level commonsense knowledge generated by a large language model (LLM) to a commonsense scene graph (CSG), supporting both types of knowledge for CSG-TL. To demonstrate the superiority of CSG-TL on target localization, extensive experiments are performed on the real-world ScanNet dataset and the AI2THOR simulator. Moreover, we have extended CSG-TL to an object search framework, CSG-OS, validated in both simulated and real-world environments. Code and videos are available at https://sites.google.com/view/csg-os. Wenqi Ge, Chao Tang 0001, Hong Zhang 0013 |
IROS | 3 |
| 2024 | SWCF-Net: Similarity-weighted Convolution and Local-global Fusion for Efficient Large-scale Point Cloud Semantic SegmentationabstractLarge-scale point cloud consists of a multitude of individual objects, thereby encompassing rich structural and underlying semantic contextual information, resulting in a challenging problem in efficiently segmenting a point cloud. Most existing researches mainly focus on capturing intricate local features without giving due consideration to global ones, thus failing to leverage semantic context. In this paper, we propose a Similarity-Weighted Convolution and local-global Fusion Network, named SWCF-Net, which takes into account both local and global features. We propose a Similarity-Weighted Convolution (SWConv) to effectively extract local features, where similarity weights are incorporated into the convolution operation to enhance the generalization capabilities. Then, we employ a downsampling operation on the K and V channels within the attention module, thereby reducing the quadratic complexity to linear, enabling Transformer to deal with large-scale point cloud. At last, orthogonal components are extracted in the global features and then aggregated with local features, thereby eliminating redundant information between local and global features and consequently promoting efficiency. We evaluate SWCF-Net on large-scale outdoor datasets SemanticKITTI and Toronto3D. Our experimental results demonstrate the effectiveness of the proposed network. Our method achieves a competitive result with less computational cost, and is able to handle large-scale point clouds efficiently. The code is available at https://github.com/Sylva-Lin/SWCF-Net. Zhenchao Lin, Li He 0002, Hongqiang Yang, Xiaoqun Sun, Guojin Zhang, Weinan Chen, Yisheng Guan, Hong Zhang 0013 |
IROS | 8 |
| 2024 | GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure DetectionabstractVisual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop closure can be fatal, so verification is required as an additional step to ensure robustness by rejecting the false positive loops. Geometric verification has been a well-acknowledged solution that leverages spatial clues provided by local feature matching to find true positives. Existing feature matching methods focus on homography and pose estimation in long-term visual localization, lacking references for geometric verification. To fill the gap, this paper proposes a unified benchmark targeting geometric verification of loop closure detection under long-term conditional variations. Furthermore, we evaluate six representative local feature matching methods (handcrafted and learning-based) under the benchmark, with in-depth analysis for limitations and future directions. Jingwen Yu, Hanjing Ye, Jianhao Jiao, Ping Tan 0002, Hong Zhang 0013 |
IROS | 5 |
| 2024 | Human Orientation Estimation Under Partial ObservationabstractReliable Human Orientation Estimation (HOE) from a monocular image is critical for autonomous agents to understand human intention. Significant progress has been made in HOE under full observation. However, the existing methods easily make a wrong prediction under partial observation and give it an unexpectedly high confidence. To solve the above problems, this study first develops a method called Part-HOE that estimates orientation from the visible joints of a target person so that it is able to handle partial observation. Subsequently, we introduce a confidence-aware orientation estimation method, enabling more accurate orientation estimation and reasonable confidence estimation under partial observation. The effectiveness of our method is validated on both public and custom-built datasets, and it shows great accuracy and reliability improvement in partial observation scenarios. In particular, we show in real experiments that our method can benefit the robustness and consistency of the Robot Person Following (RPF) task. Jieting Zhao, Hanjing Ye, Hong Zhang 0013 |
IROS | 5 |
| 2024 | Wireless Localization and Formation Control With Asynchronous AgentsabstractThe formation control of multi-agent systems has increasingly drawn attention for fulfilling numerous emerging applications and services. To achieve high-accuracy formation, the location awareness of all agents becomes an essential requirement. In this paper, we address the problem of network localization and formation control in a cooperative system with asynchronous agents. In particular, we formulate the joint localization and synchronization of agents as a statistical inference problem. The underlying probabilistic model is represented by a factor graph from which a message-passing algorithm is designed that computes approximations of the marginals of unknown variables, i.e. agents’ locations and clock offsets. Due to the Euclidean-norm operator involved in their computation no parametric closed-form expressions of the messages exist. As a compromise, implemented message-passing methods therefore resort to approximations of these messages. Conventional methods rely either on a first-order Taylor expansion of the norm operation or on non-parametric representations, e.g. by means particle filters (PFs), to compute such approximations. However, the former approach suffers from poor performance while the latter one experiences high complexity. The proposed message-passing algorithm in this paper is parametric. Specifically, it passes Gaussian messages that can be essentially obtained by suitably augmenting the factor graph and applying on it a hybrid method for combining belief propagation and variational message passing. Subsequently, the agents can exploit the estimated locations for determining the control policy. Two types of control policy are designed based on the optimization of a generalized cost function. We show that the proposed scheme enjoys a reduced complexity for multi-agent localization while achieving the desired formation with excellent accuracy. Weijie Yuan 0001, Zhaohui Yang 0001, Liangming Chen, Ruiheng Zhang 0001, Yiheng Yao, Yuanhao Cui, Hong Zhang 0013, Derrick Wing Kwan Ng |
IEEE J. Sel. Areas Commun. | 7 |
| 2024 | Hybrid Cross-Transformer-KPConv for Point Cloud SegmentationabstractPoint cloud segmentation is one of the challenging areas due to its disorder and irregularity. Currently, a lot of work utilising Transformer instead of conventional convolution methods has been proposed, which can well cope with these difficulties and is suitable for point cloud segmentation tasks. However, most existing Transformer methods extract global or local features in isolation, failing to obtain rich contextual information. In this letter, a cross-scale Transformer network for feature extration is proposed. Multi-level contextual information is captured appllying FPS algorithm. Integrated with point cloud convolution method, achieving excellent segmentation performance. Extensive experiments on SemanticKITTI dataset demonstrate the superior performance of the proposed method on mIoU. Shuhuan Wen, Pengjiang Li 0002, Hong Zhang 0013 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Robust Data Association Against Detection Deficiency for Semantic SLAMabstractRobust and accurate object association is essential for precise 3D object landmark inference in semantic Simultaneous Localization and Mapping (SLAM), and yet remains challenging due to the detection deficiency caused by high miss detection rate, false alarm, occlusion and limited field-of-view, etc. The 2D location of an object is a crucial complementary cue to the appearance feature, especially in the case of associating objects across frames under large viewpoint changes. However, motion model or trajectory pattern based methods struggle to infer object motion reliably with a moving camera. In this paper, by exploiting the local projective warping consistency, a local homography based 2D motion inference method is proposed to sequentially estimate the object location along with uncertainty. By integrating the deep appearance feature and semantic information, an object association method, named HOA, which is robust to detection deficiency is proposed. Experimental evaluations suggest that the proposed motion prediction method is capable of maintaining a low cumulative error over a long duration, which enhances the object association performance in both accuracy and robustness. Note to Practitioners—This work aims to consistently associate 2D detection boxes corresponding to the same 3D object across images. In tasks of landmark-based navigation, collision avoidance, grasping and manipulation, objects in the task space are commonly simplified into 3D enveloping surfaces (e.g. cuboid or ellipsoid) by using 2D object detection boxes from multiple image views, and accurate data association is a prerequisite for precise enveloping surface reconstruction. This problem remains challenging considering the imperfect object detections, the appearance similarity of objects and the unpredictable trajectory of the moving camera. This work proposes a long-term reliable 2D location prediction algorithm that is capable of handling the complex motion of the target. Along with the appearance feature extracted by a retrain-free deep learning based model, this work proposes an object association method that can simultaneously deal with multiple objects with unknown object categories under the moving camera scenario. Xubin Lin, Jiahao Ruan, Yirui Yang, Li He 0002, Yisheng Guan, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2024 | A Review of Cloud-Edge SLAM: Toward Asynchronous Collaboration and Implicit Representation TransmissionabstractThe utilization of cloud infrastructure and its extensive range of Internet-accessible resources holds significant potential for advancing intelligent transportation and robotics. Over the past two decades, interest in cloud-edge collaborative simultaneous localization and mapping (SLAM) has grown markedly. Consequently, a comprehensive review of current trends in this field is crucial for both novice and experienced researchers. This paper examines robots and automation systems that rely on network-based data or code, particularly in the context of SLAM development. Applying SLAM to mobile robots with limited computing power is essential for achieving autonomous navigation, and cloud-edge collaborative SLAM has emerged as an efficient solution. The review is structured around four key benefits of cloud-edge collaborative SLAM: Assisted Cloud Computing, which provides access to cloud computation and reduces the burden on edge devices; Total Cloud Computing, where the majority of computation is offloaded to the cloud, while edge devices primarily handle sensing and low-cost pre-processing; Data Storage, enabling access to large datasets, such as high-resolution environment maps and extensive training datasets, enhancing overall performance; and Data Transmission, involving cloud-edge communication for efficient data transfer and data association. Additionally, we address the challenges in existing work and the development of asynchronous collaboration and implicit representation transmission, which could mitigate transmission latency in communication-constrained environments. We believe that this review will bridge the gap between SLAM systems and deployed robotic systems, promoting the advancement of cloud-edge collaborative SLAM. Weinan Chen, Shilang Chen, Jiewu Leng, Jiankun Wang 0001, Yisheng Guan, Max Q.-H. Meng, Hong Zhang 0013 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Bridging the Gap Between Explicit and Implicit Representations: Cross-Data Association for VSLAMabstractVisual simultaneous localization and mapping (VSLAM) is a crucial technology in intelligent vehicles that relies on either explicit or implicit representations. Explicit methods are prevalent in real-time systems, offer precise geometric control, and are easy to visualize. However, they struggle with complex, dynamic environments and require high storage capacity. On the other hand, implicit techniques excel in handling intricate, changing shapes due to their compact representation and inference ability while requiring more complex display and rendering processes. A combination of both types of representations could significantly enhance the performance of VSLAM, but the cross-data association method for standalone explicit and implicit representations is still lacking. To this end, this paper proposes a data association scheme that bridges the gap between explicit and implicit representations by individually modeling the uncertainties in each representation. Our approach features a multi-level feature selection process tailored for data association. It initially extracts coarse-level features during explicit representation generation based on Bayesian estimation and refines them using the implicit representation based on ray sampling, which enhances robustness while reducing rendering costs. We rigorously evaluated our proposed methodology against current state-of-the-art approaches using public datasets and real robot scenes. The results show that our coarse-to-fine feature selection method outperforms existing techniques both quantitatively and qualitatively, suggesting its potential to significantly boost the contemporary VSLAM system performance. Shilang Chen, Xiaojie Luo, Zhenchao Lin, Shuhuan Wen, Yisheng Guan, Hong Zhang 0013, Weinan Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Visual Object Tracking With Mutual Affinity Aligned to Human IntuitionabstractSingle-object tracking generally advances by incrementally determining the tracked target's position through interactions between the search region and the template. However, the template provides less information than does the search region in terms of both temporal cues and spatial resolution. To alleviate this imbalance, we introduce an anthropic tracking framework, MATrack (Mutual Affinity Tracker), which explicitly strengthens weak template information and implicitly reduces background clutter through interactions between multiple templates and the search region. Additionally, we propose a coarse-to-fine localization approach that combines the benefits of corner-based and center-based methods. This approach enables us to simultaneously update the most recent state and background information without two-stage training. MATrack achieves state-of-the-art performance on multiple test benchmarks, including GOT-10k, LASOT, TrackingNet, OTB-100, UAV123, and NFS30. Among these benchmarks, MATrack-320's performance stands out, particularly in the short-term tracking dataset GOT-10k, where it achieves an accuracy overlap (AO) of 77.3. We also conduct comprehensive quantitative and qualitative evaluations to demonstrate that our method significantly outperforms other state-of-the-art approaches. Guotian Zeng, Bi Zeng, Qingmao Wei, Huiting Hu, Hong Zhang 0013 |
IEEE Trans. Multim. | 5 |
| 2024 | A Tightly-Coupled and Keyframe-Based Visual-Inertial-Lidar Odometry System for UGVs With Adaptive Sensor Reliability EvaluationabstractIn this article, a novel visual-inertial-lidar odometry (VILO) system named TCK-VILO is proposed to assist unmanned ground vehicles in reaching high localization performance. We introduce an adaptive sensor reliability evaluation method to configure weights of visual and lidar measurements dynamically in the backend optimization, which can improve both the accuracy and robustness of the localization in real-world outdoor scenarios. A two-stage initialization method is proposed to initialize the system by solving the spatio-temporal alignment of system states. In the TCK-VILO system, a lightweight frontend is designed to update system states while lidar or visual data is fed into the system, and then a tightly-coupled and keyframe-based backend is used to refine system states. We evaluate the proposed TCK-VILO pipeline on the public Newer College dataset and a self-collected real-world dataset. Experimental results show that our system can achieve low drifts in various challenging scenes and outperforms competing state-of-the-art VILO systems. Yan Zhuang 0013, Fei Yan 0003, Yan-Jun Liu 0003, Hong Zhang 0013 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | A Convolutional-Transformer Network for Crack Segmentation with Boundary AwarenessabstractCracks play a crucial role in assessing the safety and durability of manufactured buildings. However, the long and sharp topological features and complex background of cracks make the task of crack segmentation extremely challenging. In this paper, we propose a novel convolutional-transformer network based on encoder-decoder architecture to solve this challenge. Particularly, we designed a Dilated Residual Block (DRB) and a Boundary Awareness Module (BAM). The DRB pays attention to the local detail of cracks and adjusts the feature dimension for other blocks as needed. And the BAM learns the boundary features from the dilated crack label. Furthermore, the DRB is combined with a lightweight transformer that captures global information to serve as an effective encoder. Experimental results show that the proposed network performs better than state-of-the-art algorithms on two typical datasets. Datasets, code, and trained models are available for research at https://github.com/HqiTao/CT-crackseg. Huaqi Tao, Bingxi Liu 0001, Jinqiang Cui, Hong Zhang 0013 |
ICIP | 4 |
| 2023 | Combining Scene Coordinate Regression and Absolute Pose Regression for Visual RelocalizationabstractVisual relocalization is a fundamental problem in computer vision and robotics. Recently, regression-based methods become popular and they can be categorized into two classes: absolute pose regression and scene coordinate regression. In this work, we present a combined regression network that jointly learns scene coordinate regression and absolute pose regression for single-image visual relocalization. The proposed network composes of a feature encoder and two regression branches with uncertainty modeling. In particular, we design a deep feature conditioning module, aiming at propagating the coarse pose information in absolute pose regression to inform the predictions in scene coordinate regression. The proposed network is trained in an end-to-end fashion to learn both regression tasks. Moreover, we propose an uncertainty-driven RANSAC algorithm that incorporates the predicted scene coordinates and their uncertainties to solve the camera pose during inference. To the best of our knowledge, this work is the first to combine scene coordinate regression and pose regression in a hierarchical framework for visual relocalization. Experiments on indoor and outdoor benchmarks demonstrate the effectiveness and the superiority of the proposed method over the state-of-the-art methods. Jiahao Ruan, Li He 0002, Yisheng Guan, Hong Zhang 0013 |
ICRA | 4 |
| 2023 | Robot Person Following Under Partial OcclusionabstractRobot person following (RPF) is a capability that supports many useful human-robot-interaction (HRI) applications. However, existing solutions to person following often as-sume full observation of the tracked person. As a consequence, they cannot track the person reliably under partial occlusion where the assumption of full observation is not satisfied. In this paper, we focus on the problem of robot person following under partial occlusion caused by a limited field of view of a monocular camera. Based on the key insight that it is possible to locate the target person when one or more of hislher joints are visible, we propose a method in which each visible joint contributes a location estimate of the followed person. Experiments on a public person-following dataset show that, even under partial occlusion, the proposed method can still locate the person more reliably than the existing SOTA methods. As well, the application of our method is demonstrated in real experiments on a mobile robot. Hanjing Ye, Jieting Zhao, Yaling Pan, Weinan Chen, Li He 0002, Hong Zhang 0013 |
ICRA | 6 |
| 2023 | An Open-Source Robotic Chinese Chess PlayerabstractConsumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robot operating system, real-time and accurate chess recognition. Regarding its mechanical design, it combines a magnetism structure and mechanical cam drive, while the overall system has just three servo motors. At the same time, its control strategy is simple and effective. Furthermore, a lightweight robot message communication mechanism, entitled TinyROS, is developed for computing resource-limited embedded chips. Concerning the recognition process, our CNNbased object detector determines chess and achieves accurate identification. As a result, our robotic Chinese chess player is exquisite and easy for large-scale promotion while improving users' chess skills. Aiming to facilitate future consumer robot research and popularize customer robots, the model's mechanical and software design and the TinyROS protocol are open-sourced at https://github.com/Star-Robot/chinese-chess-robot. Shan An, Guangfu Che, Jinghao Guo, Konstantinos A. Tsintotas, Fukai Zhang, Junjie Ye 0004, Changhong Fu 0001, Haogang Zhu, Hong Zhang 0013 |
IROS | 11 |
| 2023 | Task-Oriented Grasp Prediction with Visual-Language InputsabstractTo perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a task-oriented manner (i.e., task grounding). Nevertheless, prior researches on visual-language grasping (VLG) focus on object grounding, while disregarding the fine-grained impact of tasks on object grasping. Task-incompatible grasping of a tool will inevitably limit the success of subsequent manipulation steps. Motivated by this problem, this paper proposes GraspCLIP, which addresses the challenge of task grounding in addition to object grounding to enable task-oriented grasp prediction with visual-language inputs. Evaluation on a custom dataset demonstrates that GraspCLIP achieves superior performance over established baselines with object grounding only. The effectiveness of the proposed method is further validated on an assistive robotic arm for grasping previously unseen kitchen tools given the task specification. Our presentation video is available at: https://www.youtube.com/watch?v=e1wfYQPeAXU. Chao Tang 0001, Dehao Huang, Lingxiao Meng, Hong Zhang 0013 |
IROS | 5 |
| 2023 | Multi-Scale Point Octree Encoding Network for Point Cloud Based Place RecognitionabstractOver the past decades, point cloud-based place recognition has garnered significant attention. This research paper presents a pioneering approach, denoted as the Multi-scale Point Octree Encoding Network (MPOE-Net), designed to acquire a discriminative global descriptor for efficient retrieval of places. The key element of the MPOE-Net is the point octree encoding module, which adeptly captures local information for each point by considering its nearest and farthest neighbors. Further enhancing local relationships, a multi-transformer network is introduced, utilizing a novel grouped offset-attention mechanism. To amalgamate the multi-scale attention maps into a comprehensive global descriptor, a multi-NetVLAD layer is incorporated. Through rigorous experimentation across diverse benchmark datasets, our proposed method unequivocally outperforms existing techniques in the realm of point cloud-based place recognition tasks, achieving state-of-the-art results. Our code is released publicly at https://github.com/Zhilong-Tang/MPOE-Net. Zhilong Tang, Hanjing Ye, Hong Zhang 0013 |
IROS | 3 |
| 2023 | Doubly Stochastic Distance ClusteringabstractIn doubly stochastic (DS) clustering, it is common to initialize the DS matrix with a similarity matrix and use the eigen-decomposition of the DS-scaled similarity matrix to obtain the optimal cluster indicators. The selection of a proper initial similarity measure, however, is a difficult problem and the eigen-decomposition is time-consuming, with time complexity of$O(n^{3})$where$n$is the data size. In this paper, we propose to replace the DS similarity matrix with the DS Euclidean distance matrix for clustering. We show that the optimal cluster indicators minimize the$k$-medoids error of data with DS Euclidean distance. We propose a fast method to obtain data with DS distance for clustering. Compared with DS similarity clustering, DS distance clustering is kernel-free and of low time complexity, typically$O(nd^{2}+d^{3})$where$d$is the input dimension. Experimental results on real-world datasets and the image segmentation task verify the superiority of our DS distance clustering over several competing methods. Li He 0002, Hong Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask PredictionabstractDynamic model pruning is a recent direction that allows for the inference of a different sub-network for each input sample during deployment. However, current dynamic methods rely on learning a continuous channel gating through regularization by inducing sparsity loss. This formulation introduces complexity in balancing different losses (e.g task loss, regularization loss). In addition, regularization based methods lack transparent tradeoff hyper- parameter selection to realize a computational budget. Our contribution is two-fold: 1) decoupled task and pruning losses. 2) Simple hyperparameter selection that enables FLOPs reduction estimation before training. Inspired by the Hebbian theory in Neuroscience: “neurons that fire together wire together”, we propose to predict a mask to process k filters in a layer based on the activation of its previous layer. We pose the problem as a self-supervised binary classification problem. Each mask predictor module is trained to predict if the log-likelihood for each filter in the current layer belongs to the top-k activated filters. The value k is dynamically estimated for each input based on a novel criterion using the mass of heatmaps. We show experiments on several neural architectures, such as VGG, ResNet and MobileNet on CIFAR and ImageNet datasets. On CIFAR, we reach similar accuracy to SOTA methods with 15% and 24% higher FLOPs reduction. Similarly in ImageNet, we achieve lower drop in accuracy with up to 13% improvement in FLOPs reduction. Sara Elkerdawy, Mostafa Elhoushi, Hong Zhang 0013, Nilanjan Ray |
CVPR | 3 |
| 2022 | Perspective Phase Angle Model for Polarimetric 3D Reconstruction
Guangcheng Chen, Li He 0002, Yisheng Guan, Hong Zhang 0013 |
ECCV (2) | 4 |
| 2022 | Are We Ready for Robust and Resilient SLAM? A Framework For Quantitative Characterization of SLAM DatasetsabstractReliability of SLAM systems is considered one of the critical requirements in modern autonomous systems. This directed the efforts to developing many state-of-the-art systems, creating challenging datasets, and introducing rigorous metrics to measure SLAM performance. However, the link between datasets and performance in the robustness/resilience context has rarely been explored. In order to fill this void, characterization of the operating conditions of SLAM systems is essential in order to provide an environment for quantitative measurement of robustness and resilience. In this paper, we argue that for proper evaluation of SLAM performance, the characterization of SLAM datasets serves as a critical first step. The study starts by reviewing previous efforts for quantitative characterization of SLAM datasets. Then, the problem of perturbation characterization is discussed and the linkage to SLAM robustness/resilience is established. After that, we pro-pose a novel, generic and extendable framework for quantitative analysis and comparison of SLAM datasets. Additionally, a description of different characterization parameters is provided. Finally, we demonstrate the application of our framework by presenting the characterization results of three SLAM datasets: KITTI, EuroC-MAV, and TUM-VI highlighting the level of insights achieved by the proposed framework. Islam Ali, Hong Zhang 0013 |
IROS | 2 |
| 2022 | Keyframe Selection with Information Occupancy Grid Model for Long-term Data AssociationabstractAs the basics of Visual Simultaneous Localization And Mapping (VSLAM), keyframes play an essential role. In previous works, keyframes are selected according to a series of view change-based strategies for short-term data association (STDA). However, the texture enrichment of frames is always ignored, resulting in the failure of long-term data association (LTDA). In this paper, we propose an information enrichment selection strategy with an information occupancy grid model and a deep descriptor. Frame is expressed by a deep global descriptor for a statistical explainable abstraction, in which the texture enrichment is indicated. Based on the abstraction, an information occupancy grid model is established to measure the information enrichment and the potential LTDA ability. Evaluations on variant datasets are conducted, showing the advantage of our proposed method in terms of keyframe selection and tracking precision. Also, the statistical explainability of the deep descriptor is provided. The proposed keyframe selection strategy can improve LTDA and tracking precision, especially in situations with repeated observations and loop-closures. Weinan Chen, Hanjing Ye, Chao Tang 0001, Changfei Fu, Hong Zhang 0013 |
IROS | 7 |
| 2022 | Relationship Oriented Semantic Scene Understanding for Daily Manipulation TasksabstractAssistive robot systems have been developed to help people accomplish daily manipulation tasks especially for those with disabilities, where scene understanding plays a crucial role in enabling robots to interpret the surroundings and behave accordingly. Most of the current systems approach scene understanding without considering the functional dependencies between objects. However, it is only valuable to interact with some objects when their function-relevant counterparts are considered. In this paper, we augment an assistive robotic arm system with an end-to-end semantic relationship reasoning model. It incorporates functional relationships between pairs of objects for semantic scene understanding. To ensure good generalization to unseen objects and relationships, the model works in a category-agnostic manner. We evaluate our design and three baseline methods on a self-collected benchmark with two levels of difficulty. To further demonstrate the effectiveness, the model is integrated with a symbolic planner for goal-oriented, multi-step manipulation task on a real-world assistive robotic arm platform. Chao Tang 0001, Jingwen Yu, Weinan Chen, Bingyi Xia, Hong Zhang 0013 |
IROS | 5 |
| 2022 | NDD: A 3D Point Cloud Descriptor Based on Normal Distribution for Loop Closure DetectionabstractLoop closure detection is a key technology for long-term robot navigation in complex environments. In this paper, we present a global descriptor, named Normal Distribution Descriptor (NDD), for 3D point cloud loop closure detection. The descriptor encodes both the probability density score and entropy of a point cloud as the descriptor. We also propose a fast rotation alignment process and use correlation coefficient as the similarity between descriptors. Experimental results show that our approach outperforms the state-of-the-art point cloud descriptors in both accuracy and efficency. The source code is available and can be integrated into existing LiDAR odometry and mapping (LOAM) systems. Li He 0002, Hong Zhang 0013, Xubin Lin, Yisheng Guan |
IROS | 3 |
| 2022 | Real-Time 3-D Semantic Scene Parsing With LiDAR SensorsabstractThis article proposes a novel deep-learning framework, called RSSP, for real-time 3-D scene understanding with LiDAR sensors. To this end, we introduce new sparse strided operations based on the sparse tensor representation of point clouds. Compared with conventional convolution operations, the time and space complexity of our sparse strided operations are proportional to the number of occupied voxels${N}$rather than the input spatial size${r} ^{3}$(oftenN$\ll $r3for LiDAR data). This enables our method to process point clouds at high resolutions (e.g., 20483) with a high speed (130 ms for classifying a single frame from Velodyne HDL-64). The main structure includes a CNN model built upon our sparse strided operations and a conditional random field (CRF) model to impose spatial consistency on the final predictions. A highly parallel implementation of our system is presented for both CPU-GPU and CPU-only environments. The efficiency and effectiveness of our approach are demonstrated on two public datasets (Semantic3D.net and KITTI). The experimental results and benchmark tests show that our system can be effectively applied for online 3-D data analyses with comparable or better accuracy than the state-of-the-art methods. Fei Wang 0041, Yan Zhuang 0013, Hong Zhang 0013, Hong Gu 0003 |
IEEE Trans. Cybern. | 3 |
| 2022 | Fast ORB-SLAM Without Keypoint DescriptorsabstractIndirect methods for visual SLAM are gaining popularity due to their robustness to environmental variations. ORB-SLAM2 (Mur-Artal and Tardós, 2017) is a benchmark method in this domain, however, it consumes significant time for computing descriptors that never get reused unless a frame is selected as a keyframe. To overcome these problems, we present FastORB-SLAM which is light-weight and efficient as it tracks keypoints between adjacent frames without computing descriptors. To achieve this, a two stage descriptor-independent keypoint matching method is proposed based on sparse optical flow. In the first stage, we predict initial keypoint correspondences via a simple but effective motion model and then robustly establish the correspondences via pyramid-based sparse optical flow tracking. In the second stage, we leverage the constraints of the motion smoothness and epipolar geometry to refine the correspondences. In particular, our method computes descriptors only for keyframes. We test FastORB-SLAM on TUM and ICL-NUIM RGB-D datasets and compare its accuracy and efficiency to nine existing RGB-D SLAM methods. Qualitative and quantitative results show that our method achieves state-of-the-art accuracy and is about twice as fast as the ORB-SLAM2. Qiang Fu 0013, Hongshan Yu, Xiaolong Wang 0005, Zhengeng Yang, Yong He 0012, Hong Zhang 0013, Ajmal Mian |
IEEE Trans. Image Process. | 6 |
| 2021 | Deep Snapshot Hdr Reconstruction Based On The Polarization CameraabstractThe recent development of the on-chip micro-polarizer technology has made it possible to acquire four spatially aligned and temporally synchronized polarization images with the same ease of operation as a conventional camera. In this paper, we investigate the use of this sensor technology in high-dynamic-range (HDR) imaging. Specifically, observing that natural light can be attenuated differently by varying the orientation of the polarization filter, we treat the multiple images captured by the polarization camera as a set captured under different exposure times. In our approach, we first study the relationship among polarizer orientation, degree and angle of polarization of light to the exposure time of a pixel in the polarization image. Subsequently, we propose a deep snapshot HDR reconstruction framework to recover an HDR image using the polarization images. A polarized HDR dataset is created to train and evaluate our approach. We demonstrate that our approach performs favorably against state-of-the-art HDR reconstruction algorithms. Juiwen Ting, Xuesong Wu 0001, Kangkang Hu, Hong Zhang 0013 |
ICIP | 4 |
| 2021 | Robust Improvement in 3D Object Landmark Inference for Semantic MappingabstractRecent works on semantic Simultaneous Localization and Mapping (SLAM) utilizing object landmarks have shown superiority in terms of robustness and accuracy in tracking and localization. 3D object landmarks represented by a cubic or quadric surface are inferred from 2D object bounding boxes which are typically captured from multiple views by an object detector. Nevertheless, bounding box noises and small camera baseline may lead to an inaccurate 3D object landmark inference. Inspired by the dual quadric enveloping property, in this work, we introduce the horizontal support assumption to constrain rotation w.r.t. roll and pitch for a quadric representation. As the result, we reduce the number of quadric parameters and narrow down the solution space, and ultimately produce a relatively accurate inference. Extensive experimental evaluations under both simulated and real scenarios are conducted in this paper. Quantitative results demonstrate that our approach outperforms the state-of-the-art. Xubin Lin, Yirui Yang, Li He 0002, Weinan Chen, Yisheng Guan, Hong Zhang 0013 |
ICRA | 6 |
| 2021 | UVIP: Robust UWB aided Visual-Inertial Positioning System for Complex Indoor EnvironmentsabstractIndoor positioning without GPS is a challenge task, especially, in complex scenes or when sensors fail. In this paper, we develop an ultra-wideband aided visual-inertial positioning system (UVIP) which aims to achieve accurate and robust positioning results in complex indoor environments. To this end, a point-line-based stereo visual-inertial odometry (PL-sVIO) is firstly designed to improve the positioning accuracy in structured or low-textured scenarios by making use of line features. Secondly, a loop closure method is proposed to suppress the drift of PL-sVIO based on image patch features described by a CNN for handing the situation of a large environment and viewpoint variation. Thirdly, an accurate relocalization approach is presented for the case when the visual sensor fails. In this scheme, a top-to-down matching strategy from image to point and line features is presented to improve relocalization performance. Finally, the UWB sensor is combined with the visual-inertial system to further improve the accuracy and robustness of the positioning system and provide the results in a fixed reference frame. Thus, desirable real-time positioning results are derived for complex indoor scenes. Evaluations on challenging public datasets and real-world experiments are conducted to demonstrate that the proposed UVIP can provide more accurate and robust positioning results in complex indoor environments, even in the case when the visual sensor fails or in the absence of UWB anchors. Bo Yang 0019, Jun Li 0033, Hong Zhang 0013 |
ICRA | 3 |
| 2021 | DeepRelativeFusion: Dense Monocular SLAM using Single-Image Relative Depth PredictionabstractTraditional monocular visual simultaneous localization and mapping (SLAM) algorithms have been extensively studied and proven to reliably recover a sparse structure and camera motion. Nevertheless, the sparse structure is still insufficient for scene interaction, e.g., visual navigation and augmented reality applications. To densify the scene reconstruction, the use of single-image absolute depth prediction from convolutional neural networks (CNNs) for filling in the missing structure has been proposed. However, the prediction accuracy tends to not generalize well on scenes that are different from the training datasets.In this paper, we propose a dense monocular SLAM system, named DeepRelativeFusion, that is capable to recover a globally consistent 3D structure. To this end, we use a visual SLAM algorithm to reliably recover the camera poses and semi-dense depth maps of the keyframes, and then use relative depth prediction to densify the semi-dense depth maps and refine the keyframe pose-graph. To improve the semi-dense depth maps, we propose an adaptive filtering scheme, which is a structure- preserving weighted average smoothing filter that takes into account the pixel intensity and depth of the neighbouring pixels, yielding substantial reconstruction accuracy gain in densification. To perform densification, we introduce two incremental improvements upon the energy minimization framework proposed by DeepFusion: (1) an improved cost function, and(2) the use of single-image relative depth prediction. After densification, we update the keyframes with two-view consistent optimized semi-dense and dense depth maps to improve pose- graph optimization, providing a feedback loop to refine the keyframe poses for accurate scene reconstruction. Our system outperforms the state-of-the-art dense SLAM systems quantitatively in dense reconstruction accuracy by a large margin.For more information, see the demo video and supplementary material. Shing Yan Loo, Syamsiah Mashohor, Sai Hong Tang, Hong Zhang 0013 |
IROS | 4 |
| 2021 | A multi-dimensional association information analysis approach to automated detection and localization of myocardial infarction
Jieshuo Zhang, Peng Xiong, Haiman Du, Hong Zhang 0013, Feng Lin 0002, Zeng-Guang Hou |
Eng. Appl. Artif. Intell. | 5 |
| 2021 | Learning Representations From Skeletal Self-Similarities for Cross-View Action RecognitionabstractExisting research attention in vision-based action recognition is generally paid on recognizing actions from the same views seen in the training data. One of the big challenges in action recognition lies in the large variations of action representations as actions are captured from totally different viewpoints. This paper addresses this problem by learning view-invariant representations from skeletal self-similarities of varying scales with a very light multi-stream neural network (MSNN). As human skeletons have been proved to be an effective feature modality used for action recognition and are easy to obtain, we first create a view-invariant action description by formulating skeletal self-similarities at each frame as an image (SSI), which can show a high structural stability under view changes. Accordingly, a MSNN is designed based on 3D CNN and LSTM units to learn representations from SSIs of multiple scales, where the scheme of multiple scales provides our method with a good robustness to view changes. In addition, we integrate the computation of SSIs into the MSNN by wrapping it as a custom learnable layer thanks to its simplicity, instead of normalizing and transforming skeletons using a hand-crafted preprocessing. Extensive experimental evaluations on three challenging cross-view datasets demonstrate the effectiveness of our proposed method, which achieves superior performance to the state-of-the-art algorithms on cross-view recognition. The source code of this work will be released shortly to facilitate future studies in this field. Zhanpeng Shao, Youfu Li 0001, Hong Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | To Filter Prune, or to Layer Prune, That Is the Question
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh, Hong Zhang 0013, Nilanjan Ray |
ACCV (3) | 4 |
| 2020 | One-Shot Layer-Wise Accuracy Approximation For Layer PruningabstractRecent advances in neural networks pruning have made it possible to remove a large number of filters without any perceptible drop in accuracy. However, the gain in speed depends on the number of filters per layer. In this paper, we propose a one-shot layer-wise proxy classifier to estimate layer importance that in turn allows us to prune a whole layer. In contrast to existing filter pruning methods which attempt to reduce the layer width of a dense model, our method reduces its depth and can thus guarantee inference speed up. In our proposed method, we first go through the training data once to construct proxy classifiers for each layer using imprinting. Next, we prune layers with smallest accuracy difference from their preceding layer till a latency budget is achieved. Finally, we fine-tune the newly pruned model to improve accuracy. Experimental results showed 43.70% latency reduction with 1.27% accuracy increase on CIFAR100 for the pruned VGG19. Further, we achieved 16% and 25% latency reduction with 0.58% increase and 0.01% decrease in accuracy respectively on ImageNet for ResNet-50. The major advantage of our proposed method is that these latency reductions cannot be achieved with existing filter pruning methods as they are bounded by the original model's depth. Code is available at https://github.com/selkerdawy/one-shot-layer-pruning. Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh, Hong Zhang 0013, Nilanjan Ray |
ICIP | 4 |
| 2020 | Keypoint Description by Descriptor Fusion Using AutoencodersabstractKeypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo conditional changes such as those in illumination and viewpoint. In this paper, a descriptor fusion model (DFM) is proposed to create a robust keypoint descriptor by fusing CNN-based descriptors using autoencoders. Our DFM architecture can be adapted to either trained or pre-trained CNN models. Based on the performance of existing CNN descriptors, we choose HardNet and DenseNet169 as representatives of trained and pre-trained descriptors. Our proposed DFM is evaluated on the latest benchmark datasets in computer vision with challenging conditional changes. The experimental results show that DFM is able to achieve state-of-the-art performance, with the mean mAP that is 6.45% and 6.53% higher than HardNet and DenseNet169, respectively. Zhuang Dai, Xinghong Huang, Weinan Chen, Chuangbing Chen, Li He 0002, Shuhuan Wen, Hong Zhang 0013 |
ICRA | 7 |
| 2020 | A Flexible Method for Performance Evaluation of Robot LocalizationabstractAn important research issue in mobile robotics is performance assessment of robot SLAM algorithms in terms of their localization accuracy. Typically, SLAM algorithms are evaluated with the help of benchmark datasets or expensive equipment such as motion capture. Benchmark datasets however, are environment-specific, and use of motion capture constrains spatial coverage and affordability. In this paper, we present a novel method for SLAM performance evaluation, which only uses distinctive markers (such as AR tags), randomly placed in the robot navigation environment at arbitrary locations, and observes these markers with a camera onboard of the robot. Formulated as a generative latent optimization (GLO) problem, our method uses the local robot-to-marker poses to evaluate the global robot pose estimates by a SLAM algorithm and therefore its performance. Through extensive experiments on two robots, three localization/SLAM algorithms and both LiDAR and RGB-D sensors, we demonstrate the feasibility and accuracy of our proposed method. Sean Scheideman, Nilanjan Ray, Hong Zhang 0013 |
ICRA | 3 |
| 2020 | Improving Visual Localization Accuracy in Dynamic Environments Based on Dynamic Region RemovalabstractVisual localization is a fundamental capability in robotics and has been well studied for recent decades. Although many state-of-the-art algorithms have been proposed, great success usually builds on the assumption that the working environment is static. In most of the real scenes, the assumption cannot hold because there are inevitably moving objects, especially humans, which significantly degrade the localization accuracy. To address this problem, we propose a robust visual localization system building on top of a feature-based visual simultaneous localization and mapping algorithm. We design a dynamic region detection method and use it to preprocess the input frame. The detection process is achieved in a Bayesian framework which considers both the prior knowledge generated from an object detection process and observation information. After getting the detection result, feature points extracted from only the static regions will be used for further visual localization. We performed the experiments on the public TUM data set and our recorded data set, which shows the daily dynamic scenarios. Both qualitative and quantitative results are provided to show the feasibility and effectiveness of the proposed method. Jiyu Cheng, Hong Zhang 0013, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Faster R-CNN With Classifier Fusion for Automatic Detection of Small FruitsabstractFruit detection is a fundamental task for automatic yield estimation. The goal is to detect all the fruits in images. The-state of the art of fruit detection algorithm, Faster R-CNN, shows a lack of detection advantage on small fruits. One of the reasons is only that single-level features and a classifier are used for localization of proposal candidates. In this article, we propose to incorporate a multiple classifier fusion strategy into a Faster R-CNN network for small fruit detection. We utilize features from three different levels to learn three classifiers for objectness classification in the stage of proposal localization. Probabilities from classifiers are combined by a simple convolutional layer to generate final objectness classification for proposal candidates. During training, in order to train a model with strong generalization capability, we propose to use correlation coefficients to measure the diversity of multiple classifiers. A novel loss function with classifier correlation is introduced to train the region proposal network. We evaluate the proposed model on two data sets of small fruits. Extensive experiments show that the proposed model outperforms the state-of-the-art detectors for fruit detection. Xiaochun Mai, Hong Zhang 0013, Xiao Jia 0005, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Discriminative Multi-View Privileged Information Learning for Image Re-RankingabstractConventional multi-view re-ranking methods usually perform asymmetrical matching between the region of interest (ROI) in the query image and the whole target image for similarity computation. Due to the inconsistency in the visual appearance, this practice tends to degrade the retrieval accuracy particularly when the image ROI, which is usually interpreted as the image objectness, accounts for a smaller region in the image. Since Privileged Information (PI), which can be viewed as the image prior, is able to characterize well the image objectness, we are aiming at leveraging PI for further improving the performance of multi-view re-ranking in this paper. Towards this end, we propose a discriminative multi-view re-ranking approach in which both the original global image visual contents and the local auxiliary PI features are simultaneously integrated into a unified training framework for generating the latent subspaces with sufficient discriminating power. For the on-the-fly re-ranking, since the multi-view PI features are unavailable, we only project the original multi-view image representations onto the latent subspace, and thus the re-ranking can be achieved by computing and sorting the distances from the multi-view embeddings to the separating hyperplane. Extensive experimental evaluations on the two public benchmarks, Oxford5k and Paris6k, reveal that our approach provides further performance boost for accurate image re-ranking, whilst the comparative study demonstrates the advantage of our method against other multi-view re-ranking methods. Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Hong Zhang 0013 |
IEEE Trans. Image Process. | 6 |
| 2019 | Moving Object Detection Under Discontinuous Change in Illumination Using Tensor Low-Rank and Invariant Sparse DecompositionabstractAlthough low-rank and sparse decomposition based methods have been successfully applied to the problem of moving object detection using structured sparsity-inducing norms, they are still vulnerable to significant illumination changes that arise in certain applications. We are interested in moving object detection in applications involving time-lapse image sequences for which current methods mistakenly group moving objects and illumination changes into foreground. Our method relies on the multilinear (tensor) data low-rank and sparse decomposition framework to address the weaknesses of existing methods. The key to our proposed method is to create first a set of prior maps that can characterize the changes in the image sequence due to illumination. We show that they can be detected by a k-support norm. To deal with concurrent, two types of changes, we employ two regularization terms, one for detecting moving objects and the other for accounting for illumination changes, in the tensor low-rank and sparse decomposition formulation. Through comprehensive experiments using challenging datasets, we show that our method demonstrates a remarkable ability to detect moving objects under discontinuous change in illumination, and outperforms the state-of-the-art solutions to this challenging problem. Moein Shakeri, Hong Zhang 0013 |
CVPR | 2 |
| 2019 | Lightweight Monocular Depth Estimation Model by Joint End-to-End Filter PruningabstractConvolutional neural networks (CNNs) have emerged as the state-of-the-art in multiple vision tasks including depth estimation. However, memory and computing power requirements remain as challenges to be tackled in these models. Monocular depth estimation has significant use in robotics and virtual reality that requires deployment on low-end devices. Training a small model from scratch results in a significant drop in accuracy and it does not benefit from pre-trained large models. Motivated by the literature of model pruning, we propose a lightweight monocular depth model obtained from a large trained model. This is achieved by removing the least important features with a novel joint end-to-end filter pruning. We propose to learn a binary mask for each filter to decide whether to drop the filter or not. These masks are trained jointly to exploit relations between filters at different layers as well as redundancy within the same layer. We show that we can achieve around 5x compression rate with small drop in accuracy on the KITTI driving dataset. We also show that masking can improve accuracy over the baseline with fewer parameters, even without enforcing compression loss. Sara Elkerdawy, Hong Zhang 0013, Nilanjan Ray |
ICIP | 2 |
| 2019 | A Comparison of CNN-Based and Hand-Crafted Keypoint DescriptorsabstractKeypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo condition changes such as those in illumination and viewpoint. Previous research in keypoint description has pursued three classes of descriptors: hand-crafted, those from trained convolutional neural networks (CNN), and those from pre-trained CNNs. This paper provides a comparative study of the three classes of keypoint descriptors, in terms of their ability to handle conditional changes. The study is conducted on the latest benchmark datasets in computer vision with challenging conditional changes. Our study finds that (a) in general CNN-based descriptors outperform hand-crafted descriptors, (b) the trained CNN descriptors perform better than pre-trained CNN descriptors with respect to viewpoint changes, and (c) pre-trained CNN descriptors perform better than trained CNN descriptors with respect to illumination changes. These findings can serve as a basis for selecting appropriate keypoint descriptors for various applications. Zhuang Dai, Xinghong Huang, Weinan Chen, Li He 0002, Hong Zhang 0013 |
ICRA | 5 |
| 2019 | Improving Keypoint Matching Using a Landmark-Based Image RepresentationabstractMotivated by the need to improve the performance of visual loop closure verification via multi-view geometry (MVG) under significant illumination and viewpoint changes, we propose a keypoint matching method that uses landmarks as an intermediate image representation in order to leverage the power of deep learning. In environments with various changes, the traditional verification method via MVG may encounter difficulty because of their inability to generate a sufficient number of correctly matched keypoints. Our method exploits the excellent invariance properties of convolutional neural network (ConvNet) features, which have shown outstanding performance for matching landmarks between images. By generating and matching landmarks first in the images and then matching the keypoints within the matched landmark pairs, we can significantly improve the quality of matched keypoints in terms of precision and recall measures. The proposed method is validated on challenging datasets that involve significant illumination and viewpoint changes, to establish its superior performance to the standard keypoint matching method. Xinghong Huang, Zhuang Dai, Weinan Chen, Li He 0002, Hong Zhang 0013 |
ICRA | 5 |
| 2019 | CNN-SVO: Improving the Mapping in Semi-Direct Visual Odometry Using Single-Image Depth PredictionabstractReliable feature correspondence between frames is a critical step in visual odometry (VO) and visual simultaneous localization and mapping (V-SLAM) algorithms. In comparison with existing VO and V-SLAM algorithms, semi-direct visual odometry (SVO) has two main advantages that lead to state-of-the-art frame rate camera motion estimation: direct pixel correspondence and efficient implementation of probabilistic mapping method. This paper improves the SVO mapping by initializing the mean and the variance of the depth at a feature location according to the depth prediction from a single-image depth prediction network. By significantly reducing the depth uncertainty of the initialized map point (i.e., small variance centred about the depth prediction), the benefits are twofold: reliable feature correspondence between views and fast convergence to the true depth in order to create new map points. We evaluate our method with two outdoor datasets: KITTI dataset and Oxford Robotcar dataset. The experimental results indicate that improved SVO mapping results in increased robustness and camera tracking accuracy. The implementation of this work is available at https: //github.com/yan99033/CNN-SVO Shing Yan Loo, Ali Jahani Amiri, Syamsiah Mashohor, Sai Hong Tang, Hong Zhang 0013 |
ICRA | 5 |
| 2019 | Multi-lead model-based ECG signal denoising by guided filter
Huaqing Hao, Peng Xiong, Haiman Du, Hong Zhang 0013, Feng Lin 0002, Zeng-Guang Hou |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Fast Large-Scale Spectral Clustering via Explicit Feature MappingabstractWe propose an efficient spectral clustering method for large-scale data. The main idea in our method consists of employing random Fourier features to explicitly represent data in kernel space. The complexity of spectral clustering thus is shown lower than existing Nyström approximations on largescale data. With m training points from a total of n data points, Nyström method requires O(nmd + m3+ nm2) operations, where d is the input dimension. In contrast, our proposed method requires O(nDd + D3+ n'D2), where n' is the number of data points needed until convergence and D is the kernel mapped dimension. In large-scale datasets where n ≪ n hold true, our explicitly mapping method can significantly speed up eigenvector approximation and benefit prediction speed in spectral clustering. For instance, on MNIST (60000 data points), the proposed method is similar in clustering accuracy to Nyström methods while its speed is twice as fast as Nyström. Li He 0002, Nilanjan Ray, Yisheng Guan, Hong Zhang 0013 |
IEEE Trans. Cybern. | 4 |
| 2018 | Faster R-CNN with Classifier Fusion for Small Fruit DetectionabstractThe-state-of-the-art of fruit detection with Faster R-CNN shows lack of detection advantage on small fruits. One of reasons is only single level features is used for localization of proposal candidates. In this paper, we propose to incorporate a multiple classifier fusion strategy into a Faster R-CNN network for small fruit detection. We utilize features from three different levels to learn three classifiers for objectness classification in the stage of proposal localization. Probabilities from classifiers are combined by a simple convolutional layer to generate final objectness classification for proposal candidates. In order to keep diversity of multiple classifiers, a novel loss term of classifier correlation is introduced into original loss function. Experimental results show that our model is feasible for detecting small fruits. Xiaochun Mai, Hong Zhang 0013, Max Q.-H. Meng |
ICRA | 2 |
| 2018 | Submap-Based Pose-Graph Visual SLAM: A Robust Visual Exploration and Localization System* The work in this paper is supported by the National Natural Science Foundation of China (61603103, 61673125), the Natural Science Foundation of Guangdong of China (2016A030310293), and the Major Scientific and Technological Special Project of Guangdong of China (2016B090910003)abstractFor VSLAM (Visual Simultaneous Localization and Mapping), localization is a challenging task, especially for some challenging situations: textureless frames, motion blur, etc. To build a robust exploration and localization system in a given space, a submap-based VSLAM system is proposed in this paper. Our system uses a submap back-end and a visual front-end. The main advantage of our system is its robustness with respect to tracking failure, a common problem in current VSLAM algorithms. The robustness of our system is compared with the state-of-the-art in terms of average tracking percentage. The precision of our system is also evaluated in terms of ATE (absolute trajectory error) RMSE (root mean square error) comparing the state-of-the-art. The ability of our system in solving the “kidnapped” problem is demonstrated. Our system can improve the robustness of visual localization in challenging situations. Weinan Chen, Yisheng Guan, C. Ronald Kube, Hong Zhang 0013 |
IROS | 5 |
| 2018 | Fast Shadow Detection from a Single Image Using a Patched Convolutional Neural NetworkabstractIn recent years, various shadow detection methods from a single image have been proposed and used in vision systems; however, most of them are not appropriate for the robotic applications due to the expensive time complexity. This paper introduces a fast shadow detection method using a deep learning framework, with a time cost that is appropriate for robotic applications. In our solution, we first obtain a shadow prior map with the help of multi-class support vector machine using statistical features. Then, we use a semantic-aware patch-level Convolutional Neural Network that efficiently trains on shadow examples by combining the original image and the shadow prior map. Experiments on benchmark datasets demonstrate the proposed method significantly decreases the time complexity of shadow detection, by one or two orders of magnitude compared with state-of-the-art methods, without losing accuracy. Sepideh Hosseinzadeh, Moein Shakeri, Hong Zhang 0013 |
IROS | 3 |
| 2018 | Real-Time Segmentation with Appearance, Motion and GeometryabstractReal-time Segmentation is of crucial importance to robotics related applications such as autonomous driving, driving assisted systems, and traffic monitoring from unmanned aerial vehicles imagery. We propose a novel two-stream convolutional network for motion segmentation, which exploits flow and geometric cues to balance the accuracy and computational efficiency trade-offs. The geometric cues take advantage of the domain knowledge of the application. In case of mostly planar scenes from high altitude unmanned aerial vehicles (UAVs), homography compensated flow is used. While in the case of urban scenes in autonomous driving, with GPS/IMU sensory data available, sparse projected depth estimates and odometry information are used. The network provides 4.7× speedup over the state of the art networks in motion segmentation from 153ms to 36ms, at the expense of a reduction in the segmentation accuracy in terms of pixel boundaries. This enables the network to perform real-time on a Jetson T×2. In order to recuperate some of the accuracy loss, geometric priors is used while still achieving a much improved computational efficiency with respect to the state-of-the-art. The usage of geometric priors improved the segmentation in UAV imagery by 5.2 % using the metric of IoU over the baseline network. While on KITTI-MoSeg the sparse depth estimates improved the segmentation by 12.5 % over the baseline. Our proposed motion segmentation solution is verified on the popular KITTI and VIVID datasets, with additional labels we have produced. The code for our work is publicly available at1. Mennatullah Siam, Sara Elkerdawy, Mostafa Gamal, Moemen Abdel-Razek, Martin Jägersand, Hong Zhang 0013 |
IROS | 6 |
| 2018 | Object Classification With Joint Projection and Low-Rank Dictionary LearningabstractFor an object classification system, the most critical obstacles toward real-world applications are often caused by large intra-class variability, arising from different lightings, occlusion, and corruption, in limited sample sets. Most methods in the literature would fail when the training samples are heavily occluded, corrupted or have significant illumination or viewpoint variations. Besides, most of the existing methods and especially deep learning-based methods, need large training sets to achieve a satisfactory recognition performance. Although using the pre-trained network on a generic large-scale data set and fine-tune it to the small-sized target data set is a widely used technique, this would not help when the content of base and target data sets are very different. To address these issues simultaneously, we propose a joint projection and low-rank dictionary learning method using dual graph constraints. Specifically, a structured class-specific dictionary is learned in the low-dimensional space, and the discrimination is further improved by imposing a graph constraint on the coding coefficients, that maximizes the intra-class compactness and inter-class separability. We enforce structural incoherence and low-rank constraints on sub-dictionaries to reduce the redundancy among them, and also make them robust to variations and outliers. To preserve the intrinsic structure of data, we introduce a supervised neighborhood graph into the framework to make the proposed method robust to small-sized and high-dimensional data sets. Experimental results on several benchmark data sets verify the superior performance of our method for object classification of small-sized data sets, which include a considerable amount of different kinds of variation, and may have high-dimensional feature vectors. Homa Foroughi, Nilanjan Ray, Hong Zhang 0013 |
IEEE Trans. Image Process. | 3 |
| 2018 | Kernel K-Means Sampling for Nyström ApproximationabstractA fundamental problem in Nyström-based kernel matrix approximation is the sampling method by which training set is built. In this paper, we suggest to use kernel -means sampling, which is shown in our works to minimize the upper bound of a matrix approximation error. We first propose a unified kernel matrix approximation framework, which is able to describe most existing Nyström approximations under many popular kernels, including Gaussian kernel and polynomial kernel. We then show that, the matrix approximation error upper bound, in terms of the Frobenius norm, is equal to the -means error of data points in kernel space plus a constant. Thus, the -means centers of data in kernel space, or the kernel -means centers, are the optimal representative points with respect to the Frobenius norm error upper bound. Experimental results, with both Gaussian kernel and polynomial kernel, on real-world data sets and image segmentation tasks show the superiority of the proposed method over the state-of-the-art methods. Li He 0002, Hong Zhang 0013 |
IEEE Trans. Image Process. | 2 |
| 2017 | Moving Object Detection in Time-Lapse or Motion Trigger Image Sequences Using Low-Rank and Invariant Sparse DecompositionabstractLow-rank and sparse representation based methods have attracted wide attention in background subtraction and moving object detection, where moving objects in the scene are modeled as pixel-wise sparse outliers. Since in real scenarios moving objects are also structurally sparse, recently researchers have attempted to extract moving objects using structured sparse outliers. Although existing methods with structured sparsity-inducing norms produce promising results, they are still vulnerable to various illumination changes that frequently occur in real environments, specifically for time-lapse image sequences where assumptions about sparsity between images such as group sparsity are not valid. In this paper, we first introduce a prior map obtained by illumination invariant representation of images. Next, we propose a low-rank and invariant sparse decomposition using the prior map to detect moving objects under significant illumination changes. Experiments on challenging benchmark datasets demonstrate the superior performance of our proposed method under complex illumination changes. Moein Shakeri, Hong Zhang 0013 |
ICCV | 2 |
| 2017 | Face recognition using multi-modal low-rank dictionary learningabstractFace recognition has been widely studied due to its importance in different applications; however, most of the proposed methods fail when face images are occluded or captured under illumination and pose variations. Recently several low-rank dictionary learning methods have been proposed and achieved promising results for noisy observations. While these methods are mostly developed for single-modality scenarios, recent studies demonstrated the advantages of feature fusion from multiple inputs. We propose a multi-modal structured low-rank dictionary learning method for robust face recognition, using raw pixels of face images and their illumination invariant representation. The proposed method learns robust and discriminative representations from contaminated face images, even if there are few training samples with large intra-class variations. Extensive experiments on different datasets validate the superior performance and robustness of our method to severe illumination variations and occlusion. Homa Foroughi, Moein Shakeri, Nilanjan Ray, Hong Zhang 0013 |
ICIP | 4 |
| 2017 | Fast-SeqSLAM: A fast appearance based place recognition algorithmabstractLoop closure detection or place recognition is a fundamental problem in robot simultaneous localization and mapping (SLAM). SeqSLAM is considered to be one of the most successful algorithms for loop closure detection as it has been demonstrated to be able to handle significant environmental condition changes including those due to illumination, weather, and time of the day. However, SeqSLAM relies heavily on exhaustive sequence matching, a computationally expensive process that prevents the algorithm from being used in dealing with large maps. In this paper, we propose Fast-SeqSLAM, an efficient version of SeqSLAM. Fast-SeqSLAM has a much reduced time complexity without degrading the accuracy, and this is achieved by using an approximate nearest neighbor (ANN) algorithm to match the current image with those in the robot map and extending the idea of SeqSLAM to greedily search a sequence of images that best match with the current sequence. We demonstrate the effectiveness of our Fast-SeqSLAM algorithm in appearance based loop closure detection. Sayem Mohammad Siam, Hong Zhang 0013 |
ICRA | 2 |
| 2017 | Error bound of Nyström-approximated NCut eigenvectors and its application to training size selection
Li He 0002, Nilanjan Ray, Hong Zhang 0013 |
Neurocomputing | 3 |
| 2017 | A chordiogram image descriptor using local edgels
Xiaolong Wang 0005, Hong Zhang 0013, Guohua Peng |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | M2DP: A novel 3D point cloud descriptor and its application in loop closure detectionabstractIn this paper, we present a novel global descriptor M2DP for 3D point clouds, and apply it to the problem of loop closure detection. In M2DP, we project a 3D point cloud to multiple 2D planes and generate a density signature for points for each of the planes. We then use the left and right singular vectors of these signatures as the descriptor of the 3D point cloud. Our experimental results show that the proposed algorithm outperforms state-of-the-art global 3D descriptors in both accuracy and efficiency. Li He 0002, Xiaolong Wang 0005, Hong Zhang 0013 |
IROS | 3 |
| 2016 | Illumination invariant representation of natural images for visual place recognitionabstractIllumination changes are a typical problem for many outdoor long-term applications such as visual place recognition. Keypoints may fail to match between images taken at the same location but different times of the day. Although recently some methods are presented for creating shadow-free image representations, all of them have the limitation in terms of dealing with night images and non-Planckian source of lighting. In this paper we present a new method for creating illumination invariant image representation using a combination of two existing methods based on natural image statistics that address the issue of illumination invariance. Unlike previous attempts at solving the problem of illumination invariant representation, the proposed method does not assume the ideal narrow-band color camera nor a calibration step for each environment. We evaluate our method on real datasets to establish its accuracy and efficiency. Experimental results show that our method outperforms competing methods for illumination invariant image representation. Moein Shakeri, Hong Zhang 0013 |
IROS | 2 |
| 2016 | COROLA: A sequential solution to moving object detection using low-rank approximation
Moein Shakeri, Hong Zhang 0013 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Visual saliency detection based on multi-scale and multi-channel mean
Lang Sun, Hong Zhang 0013 |
Multim. Tools Appl. | 3 |
| 2016 | Iterative ensemble normalized cuts
Li He 0002, Hong Zhang 0013 |
Pattern Recognit. | 2 |
| 2015 | Joint Feature Selection with Low-rank Dictionary Learning
Homa Foroughi, Moein Shakeri, Nilanjan Ray, Hong Zhang 0013 |
BMVC | 4 |
| 2015 | Person tracking and following with 2D laser scannersabstractHaving accurate knowledge of the positions of people around a robot provides rich, objective and quantitative data that can be highly useful for a wide range of tasks, including autonomous person following. The primary objective of this research is to promote the development of robust, repeatable and transferable software for robots that can automatically detect, track and follow people in their environment. The work is strongly motivated by the need for such functionality onboard an intelligent power wheelchair robot designed to assist people with mobility impairments. In this paper we propose a new algorithm for robust detection, tracking and following from laser data. We show that the approach is effective in various environments, both indoor and outdoor, and on different robot platforms (the intelligent power wheelchair and a Clearpath Husky). The method has been implemented in the Robot Operating System (ROS) framework and will be publicly released as a ROS package. We also describe and will release several datasets designed to promote the standardized evaluation of similar algorithms. Angus Leigh, Joelle Pineau, Nicolas A. Olmedo, Hong Zhang 0013 |
ICRA | 4 |
| 2015 | Keypoint matching by outlier pruning with consensus constraintabstractA simple and reliable keypoint matching method is proposed in this paper. Our research is motivated by the desire to improve the performance of multi-view geometry (MVG) based verification in visual loop closure detection under significant illumination change, where traditional methods may fail due to their inability to either find a sufficient number of correctly matched keypoints or identify correct underlying camera motion to verify the matches. Our method is inspired by research on the spatial statistics of optical flow. By observing that the displacement of matching keypoints between a pair of images is equivalent to the optical flow under the assumption of small camera motion (which is true in applications such as loop closure detection), we exploit the fact that the displacement of correctly matched keypoints between two images must follow a well-defined distribution. This paves the way to a keypoint matching method that uses this distribution to screen or prune potential matching keypoints, so as to remove the incorrect matches (outliers) and retain the true matches (inliers) without being overly and solely dependent on keypoint descriptors. The proposed method is validated on the outdoor image sequences and shows superior performance to the standard keypoint matching method based on distance ratio test. Yang Liu 0272, Rong Feng, Hong Zhang 0013 |
ICRA | 3 |
| 2015 | 3-DOF point cloud registration using congruent trianglesabstractIn this paper, we present an efficient 3-DOF registration method to align two overlapping point clouds captured by a range sensor that experiences locally planar motion such as in many robotics applications. The algorithm follows the RANSAC framework and is based on finding a pair of congruent triangles in the source and target clouds. With the assumption of planar sensor motion, corresponding vertices of the congruent triangles must have similar elevation values, allowing our algorithm to identify them efficiently. Specifically, given a triangle in the source cloud, our algorithm first finds two vertices of a candidate congruent triangle in the target cloud and then uses the third vertex to verify, minimizing the complexity of the geometric base as well as the expected time of sampling successfully a matching congruent triangle. To improve the performance of our algorithm further, our verification of a hypothetical alignment transformation proceeds first locally by using only the points near the vertices of the congruent triangles before involving the entire cloud. Our experimental results show that the proposed algorithm outperforms state-of-the-art registration algorithms in the case of 3-DOF planar sensor motion. Xiaolong Wang 0005, Hong Zhang 0013, Guohua Peng |
IROS | 2 |
| 2015 | Robust people counting using sparse representation and random projection
Homa Foroughi, Nilanjan Ray, Hong Zhang 0013 |
Pattern Recognit. | 3 |
| 2014 | Computer vision and applications in Canadian oil sandsabstractCanada holds oil reserve at 168 billion barrels - in the form of oil sands in the Province of Alberta-ranking it the third largest in the world. In this lecture, I will highlight our collaborative research with Canadian oil sands producers - in computer vision, image processing and machine learning - to address some of the challenges facing this important Canadian industry. Specifically I will provide examples of how we have been able to tackle practical industrial problems successfully and, at the same time, contribute to the progress of scientific disciplines such as computer vision and image processing. These examples include: delivering performance indicators for the optimization of a mining process, monitoring the conditions of mining equipment, and protecting wildlife in and around mine sites. Our experience demonstrates that academic research can be carried out effectively in close collaboration with industry, and that such a partnership can be mutually beneficial. Hong Zhang 0013 |
ICARCV | 1 |
| 2014 | People counting with image retrieval using compressed sensingabstractThe estimation of the number of people present in an image has many applications such as intelligent transportation, urban planning and crowd surveillance. Rather than conventional counting by detection or regression/machine-learning methods, we propose an image retrieval approach, which uses an image descriptor to estimate the people count. We review the performance of several image descriptors. In addition, we propose a straightforward global image descriptor for image retrieval based on compressed sensing theory. Extensive evaluations on existing crowd analysis benchmark datasets demonstrate the effectiveness of our image retrieval-based approach compared to state-of-the-art regression-based people counting methods. Homa Foroughi, Nilanjan Ray, Hong Zhang 0013 |
ICASSP | 3 |
| 2014 | VFCCV snake: A novel active contour model combining edge and regional informationabstractActive contour models have been widely used for image segmentation. Among leading models of active contour is vector-field convolution (VFC), a parametric active contour that improves the popular gradient vector flow (GVF) model. However VFC is still sensitive to noise and can be easily trapped in cluttered regions of an image because it only considers edge information. Based on the geometric active contour model proposed by Chan and Vese, this paper introduces a novel active contour model that incorporates region information in VFC in order to take advantage of edge and regional information. This new model, which we refer to as VFCCV snake, is implemented in the parametric active contour framework, and has control on topology especially in noisy images and images with boundary gaps. Experimental results on both synthetic and real images show superior performance of our VFCCV snake to state-of-the-art leading active contour methods. Jiuyu Sun, Nilanjan Ray, Hong Zhang 0013 |
ICIP | 3 |
| 2014 | An efficient index for visual search in appearance-based SLAMabstractVector-quantization can be a computationally expensive step in visual bag-of-words (BoW) search when the vocabulary is large. A BoW-based appearance SLAM needs to tackle this problem for an efficient real-time operation. We propose an effective method to speed up the vector quantization process in BoW-based visual SLAM. We employ a graph-based nearest neighbor search (GNNS) algorithm to this aim, and experimentally show that it can outperform the state-of-the-art. The graph-based search structure used in GNNS can efficiently be integrated into the BoW model and the SLAM framework. The graph-based index, which is a k-NN graph, is built over the vocabulary words and can be extracted from the BoW's vocabulary construction procedure, by adding one iteration to the k-means clustering, which adds small extra cost. Moreover, exploiting the fact that images acquired for appearance-based SLAM are sequential, GNNS search can be initiated judiciously which helps increase the speedup of the quantization process considerably. Kiana Hajebi, Hong Zhang 0013 |
ICRA | 2 |
| 2014 | An efficient visual loop closure detection method in a map of 20 million key locationsabstractAn important problem in robot simultaneous localization and mapping (SLAM) is loop closure detection. Recent studies of the problem have led to successful development of methods that are based on images captured by the robot. These methods tackle the issue of efficiency through data structures such as indexing and hierarchical (tree) organization of the image data that represent the robot map. In this paper, we offer an alternative approach and present a novel method for visual loop-closure detection. Our approach uses an extremely simple image representation, namely, a down-sampled binarized version of the original image, combined with a highly efficient image similarity measure - mutual information. As a result, our method is able to perform loop closure detection in a map with 20 million key locations in about 2.38 seconds on a commodity computer. The excellent performance of our method in terms of its low complexity and accuracy in experiments establishes it as a promising solution to loop closure detection in large-scale robot maps. Hong Zhang 0013, Yisheng Guan |
ICRA | 2 |
| 2014 | Detection of small moving objects using a moving cameraabstractIn recent years, various background subtraction methods have been proposed and used in vision systems for moving object detection and tracking from moving cameras; however, most of them have difficulty in handling small and distant objects in complicated non-flat scenes. This paper presents a robust method to effectively segment moving objects from videos, captured by a camera on a moving platform. In our approach, a two-level registration is applied to estimate the effect of camera motion for motion compensation. After motion estimation and extraction of potential foreground pixels by Gaussian mixture model, noisy result is refined using component based and pixel based methods the latter of which uses the hidden markov model (HMM) for classifying pixels. Finally, foreground objects are tracked by a particle filter to exploit the temporal coherence of foreground motion and improve the detection accuracy through time. Experimental results show that our method outperforms competing methods for detecting moving objects in complex environments. Moein Shakeri, Hong Zhang 0013 |
IROS | 2 |
| 2014 | Convex-Relaxed Kernel Mapping for Image SegmentationabstractThis paper investigates a convex-relaxed kernel mapping formulation of image segmentation. We optimize, under some partition constraints, a functional containing two characteristic terms: 1) a data term, which maps the observation space to a higher (possibly infinite) dimensional feature space via a kernel function, thereby evaluating nonlinear distances between the observations and segments parameters and 2) a total-variation term, which favors smooth segment surfaces (or boundaries). The algorithm iterates two steps: 1) a convex-relaxation optimization with respect to the segments by solving an equivalent constrained problem via the augmented Lagrange multiplier method and 2) a convergent fixed-point optimization with respect to the segments parameters. The proposed algorithm can bear with a variety of image types without the need for complex and application-specific statistical modeling, while having the computational benefits of convex relaxation. Our solution is amenable to parallelized implementations on graphics processing units (GPUs) and extends easily to high dimensions. We evaluated the proposed algorithm with several sets of comprehensive experiments and comparisons, including: 1) computational evaluations over 3D medical-imaging examples and high-resolution large-size color photographs, which demonstrate that a parallelized implementation of the proposed method run on a GPU can bring a significant speed-up and 2) accuracy evaluations against five state-of-the-art methods over the Berkeley color-image database and a multimodel synthetic data set, which demonstrates competitive performances of the algorithm. Mohamed Ben Salah, Ismail Ben Ayed, Jing Yuan 0001, Hong Zhang 0013 |
IEEE Trans. Image Process. | 4 |
| 2014 | Object joint detection and tracking using adaptive multiple motion models
Zhijie Wang 0003, Mohamed Ben Salah, Hong Zhang 0013 |
Vis. Comput. | 3 |
| 2013 | Adaptive shape prior in graph cut image segmentation
Hong Zhang 0013, Nilanjan Ray |
Pattern Recognit. | 2 |
| 2012 | Wavelet subband-based steam detection by multiple kernel learningabstractWavelet transform coefficients have been shown as significant features for detecting steam and smoke. Wavelet transform is multi-resolution in nature; moreover, at each resolution, wavelet transform coefficients form a high dimensional feature set. In this paper we handle both these issues in a multiple kernel learning (MKL) framework. First, high dimensionality is handled by using a kernel function that measures similarity between two sets of wavelet coefficients at the same resolution. Next, we consider a convex combination of these kernel functions that correspond to all the available resolutions of the wavelet transform. The proposed MKL uses an L1norm linear support vector machine (SVM) for sparse learning of the convex combination. Then, this mixture kernel function is used in an L2norm nonlinear SVM for binary classification- image with steam or without steam. Our method yields encouraging results and outperforms other competing methods. Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013 |
ICIP | 3 |
| 2012 | Convex relaxation for image segmentation by kernel mappingabstractThis study proposes a novel multiregion image segmentation method using convex relaxation optimization and kernel mapping of the image data. The image data is transformed by a kernel function in order to support various image models while avoiding complex modeling. This is embedded implicitly in an objective function which is optimized by iterating a two-step strategy. First, a fixed point sequence is used to evaluate the regions parameters. Second, the image partition is updated by an efficient multiplier-based algorithm which uses the standard augmented Lagrangian method. A thorough experimental study is carried out over a multi-model synthetic dataset, the Berkeley database, as well as cardiac 3D data to show the effectiveness of the proposed method. Mohamed Ben Salah, Ismail Ben Ayed, Jing Yuan 0001, Zhijie Wang 0003, Hong Zhang 0013 |
ICIP | 5 |
| 2012 | Seeing through clutter: Snake computation with dynamic programming for particle segmentation
Nilanjan Ray, Scott T. Acton, Hong Zhang 0013 |
ICPR | 3 |
| 2012 | Indexing visual features: Real-time loop closure detection using a tree structureabstractWe propose a simple and effective method for visual loop closure detection in appearance-based robot SLAM. Unlike the Bag-of-Words (BoW hereafter) approach in most existing work of the problem, our method uses direct feature matching to detect loop closures and therefore avoid the perceptual aliasing problem caused by the vector quantization process of BoW. We show that a tree structure can be efficient in online loop closure detection. In our method, a KD-tree is built over all the key frame features and an indexing table is kept for retrieving relevant key frames. Due to the efficiency of the tree-based feature matching, loop closure detection can be achieved in real-time. To investigate the scalability of the method, we also apply the scale dependent feature selection in our method and show that the run time can be reduced significantly at the expense of sacrificing the performance to some extent. The proposed method is validated on an indoor SLAM dataset with 7,420 images. Yang Liu 0272, Hong Zhang 0013 |
ICRA | 2 |
| 2012 | Visual loop closure detection with a compact image descriptorabstractIn this paper, we present a method for visual loop closure detection using a compact image descriptor, Gabor-Gist. In contrast to the Bag-of-Words (BoW) approach, which is dominant in recent studies of the loop closure detection problem that derives an image descriptor from locally extracted keypoint descriptors, our method relies on a single efficient image descriptor of low dimension to describe and measure similarities among images. We employ PCA to transform a high dimensional Gabor-Gist descriptor to a lower dimensional form to improve both the computational efficiency of our method and the discriminative power of the image descriptor. In addition, we use a particle filter to exploit the correlation among images in a sequence captured by the robot in the process of identifying loop closure candidates. Our method is highly scalable due to the compactness of the image descriptor and the simplicity of particle filtering. To validate our method, we used the Oxford City dataset. Our experimental results show that for this dataset, high recall (up to 87%) can be obtained at 100% precision, with only a few particles. Yang Liu 0272, Hong Zhang 0013 |
IROS | 2 |
| 2012 | Shape based appearance model for kernel tracking
Zhijie Wang 0003, Mohamed Ben Salah, Hong Zhang 0013, Nilanjan Ray |
Image Vis. Comput. | 3 |
| 2012 | Biologically inspired collective construction with visual landmarksabstractWe describe our research in using environmental visual landmarks as the basis for completing simple robot construction tasks. Inspired by honeybee visual navigation behavior, a visual template mechanism is proposed in which a natural landmark serves as a visual reference or template for distance determination as well as for navigation during collective construction. To validate our proposed mechanism, a wall construction problem is investigated and a minimalist solution is given. Experimental results show that, using the mechanism of a visual template, a collective robotic system can successfully build the desired structure in a decentralized fashion using only local sensing and no direct communication. In addition, a particular variable, which defines tolerance for alignment of the structure, is found to impact the system performance. By decreasing the value of the variable, system performance is improved at the expense of a longer construction time. The visual template mechanism is appealing in that it can use a reference point or salient object in a natural environment that is new or unexplored and it could be adapted to facilitate more complicated building tasks. Zhengwei Zhang, Hong Zhang 0013, Yibin Li 0001 |
J. Zhejiang Univ. Sci. C | 2 |
| 2012 | Clump splitting via bottleneck detection and shape classification
Hong Zhang 0013, Nilanjan Ray |
Pattern Recognit. | 2 |
| 2012 | Shape based local thresholding for binarization of document images
Jichuan Shi, Nilanjan Ray, Hong Zhang 0013 |
Pattern Recognit. Lett. | 3 |
| 2012 | Object Detection With DoG Scale-Space: A Multiple Kernel Learning ApproachabstractDifference of Gaussians (DoG) scale-space for an image is a significant way to generate features for object detection and classification. While applying DoG scale-space features for object detection/classification, we face two inevitable issues: dealing with high dimensional data and selecting/weighting of proper scales. The scale selection process is mostly ad-hoc to date. In this paper, we propose a multiple kernel learning (MKL) method for both DoG scale selection/weighting and dealing with high dimensional scale-space data. We design a novel shift invariant kernel function for DoG scale-space. To select only the useful scales in the DoG scale-space, a novel framework of MKL is also proposed. We utilize a 1-norm support vector machine (SVM) in the MKL optimization problem for sparse weighting of scales from DoG scale-space. The optimized data-dependent kernel accommodates only a few scales that are most discriminatory according to the large margin principle. With a 2-norm SVM this learned kernel is applied to a challenging detection problem in oil sand mining: to detect large lumps in oil sand videos. We tested our method on several challenging oil sand data sets. Our method yields encouraging results on these difficult-to-process images and compares favorably against other popular multiple kernel methods. Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013 |
IEEE Trans. Image Process. | 3 |
| 2011 | Clump splitting via bottleneck detectionabstractUnder-segmentation of an image with multiple objects is a common problem in image segmentation algorithms. This paper presents a novel approach for the splitting of clumps formed by multiple objects due to under-segmentation. The algorithm includes two steps: finding a pair of points for clump splitting, and joining the pair of selected points. In the first step, a pair of points for splitting is detected using a bottleneck rule, under the assumption that the desired objects have roughly convex shape. In the second step, the selected pair of splitting points is joined by finding the optimal splitting line between them, based on minimizing an image energy. The performance of this method is evaluated using images from various applications. Experimental results show that the proposed approach has several advantages over existing splitting methods in identifying points for splitting as well as finding an accurate split line. Hong Zhang 0013, Nilanjan Ray |
ICIP | 2 |
| 2011 | A novel 6-DoF biped active walking robot - Walking gaits, patterns and experimentsabstractCombining the advantages of active and passive walking robots, we have developed a novel active biped walking robot with only six DoFs. The robot is built with six 1-DoF joint modules and two wheels as the feet. It achieves locomotion in special gaits different from those of traditional biped robots. In this paper, this novel biped robot is introduced, and four walking gaits, namely turning-around gait, foot-wheel hybrid gait, side-stepping gait and turning-over gait are proposed, and their walking patterns and motion planning are presented and analyzed. Walking experiments are carried out to verify the locomotion function, the effectiveness of the presented gaits and to illustrate the features of this novel biped robot. It has been shown that biped active walking may be achieved with only a few DoFs and simple kinematic configuration. Yisheng Guan, Xuefeng Zhou, Haifei Zhu, Chuanwu Cai, Hong Zhang 0013 |
ICRA | 7 |
| 2011 | BoRF: Loop-closure detection with scale invariant visual featuresabstractIn this paper, we present a novel method for visual loop-closure detection in autonomous robot navigation. Our method, which we refer to as bag-of-raw-features or BoRF, uses scale-invariant visual features (such as SIFT) directly, rather than their vector-quantized representation or bag-of-words (BoW), which is popular in recent studies of the problem. BoRF avoids the offline process of vocabulary construction, and does not suffer from the perceptual aliasing problem of BoW, thereby significantly improving the recall performance. To reduce the computational cost of direct feature matching, we exploit the fact that images in the case of robot navigation are acquired sequentially, and that feature matching repeatability with respect to scale can be learned and used to reduce the number of the features considered for matching. The proposed method is tested experimentally using indoor visual SLAM image sequences. Hong Zhang 0013 |
ICRA | 1 |
| 2011 | Fast Approximate Nearest-Neighbor Search with k-Nearest Neighbor Graph
Kiana Hajebi, Yasin Abbasi-Yadkori, Hossein Shahbazi, Hong Zhang 0013 |
IJCAI | 4 |
| 2011 | Climbot: A modular bio-inspired biped climbing robotabstractHigh-rise tasks in agriculture, forestry and building industry requires robots possessing climbing function. Motivated by these potential applications and inspired by the climbing motion of animals such as inchworms, we have developed a novel biped climbing robot - Climbot. Built with a modular approach, the robot consists of five 1-DoF joint modules connected in series and two special grippers mounted at the ends. With this configuration, Climbot is able not only to climb a variety of media, but also to grasp and manipulate objects, and hence is a ¿mobile¿ manipulator. In this paper, we first introduce the development of this novel robot, and then illustrate three climbing gaits based on the unique configuration of the robot. Experiments of climbing poles are carried out to verify the climbing functions and to demonstrate potential application of the proposed robot. Yisheng Guan, Haifei Zhu, Xuefeng Zhou, Chuanwu Cai, Wenqiang Wu, Zhanchu Li, Hong Zhang 0013 |
IROS | 8 |
| 2011 | Application of locality sensitive hashing to realtime loop closure detectionabstractIn this work we present a new approach for detecting loop closures in a real-time online setting. The Loop Closure Detection problem is important in visual SLAM applications and different approaches exist to deal with this problem. Most of these approaches are based on the Bag-of-Words approach, and assume a fixed visual vocabulary can work in different types of environments. However BOW is known to introduce perceptual aliasing. By using Locality Sensitive Hashing (LSH) we are able to compute image similarity and detect loop closures by using visual features directly without vector quantization as in BOW and also LSH does not require a prior visual vocabulary. We show the effectiveness of our approach empirically by comparing it to the Bag of Words (BOW) approach which is the dominant method of selecting candidate loop closing images. Our method is fast enough for realtime applications and its accuracy is significantly better than the BOW approach. Hossein Shahbazi, Hong Zhang 0013 |
IROS | 2 |
| 2011 | Exploiting SLAM to improve feature matchingabstractOne of the main problems associated with vision-based SLAM is the difficulty to find high-quality and indistinguishable landmarks which are detectable across significant viewpoint changes. Most of the current visual SLAM systems utilize the feature extraction methods that have been proposed for the offline computer vision tasks, which are either too slow for a real-time SLAM or not robust enough against viewpoint changes. In this paper, we demonstrate that by incorporating the position information of the robot into the feature description, we can help the feature matching process significantly. Kiana Hajebi, Hong Zhang 0013 |
SMC | 2 |
| 2010 | Automating Snakes for Multiple Objects Detection
Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013 |
ACCV (3) | 3 |
| 2010 | Designing compact Gabor filter banks for efficient texture feature extractionabstractTexture feature has been widely used in image segmentation, classification, retrieval and many others. Among various approaches to texture feature extraction, Gabor filtering has emerged as one of the most popular in recent years. Gabor filter-based texture feature extractor is in fact a Gabor filter bank defined by its parameters including frequencies, orientations and smoothing parameters of the Gaussian envelope. In the literature, these parameters are often set by trial and error, based on the experience of the user, and the Gabor filter banks thus designed are often over-sized. To address the problem mentioned above, we propose to design compact Gabor filter banks by incorporating filter selection in this study. We develop a new Mahalanobis separability measure-based supervised approach to address the need of texture feature extraction. The strengths of our methods are twofold. Firstly, the proposed method provides a systematic way for Gabor filter bank design to avoid man-made bias. Secondly, the compact filter banks thus designed overcomes the problem of redundant or insignificant/irrelevant filter banks, and this in turn leads to improved performance of texture classification. Experimental results on benchmark datasets demonstrate the effectiveness of our proposed approach. Kezhi Mao, Hong Zhang 0013, Tianyou Chai |
ICARCV | 3 |
| 2010 | Improving image segmentation via shape PCA reconstructionabstractThis paper proposes a post-processing method for image segmentation to take advantage of information not directly available from the image. Specifically, the proposed method improves the segmentation of an image by making use of shape information learned from training shapes in ground truth images. To obtain shape prior, training shapes are first aligned by congealing, and then landmark interpolation is performed, followed by shape PCA on aligned shapes. To improve a segmentation, subsequently, shape PCA reconstruction is performed using the first few principal components on objects in the segmented image. Shape PCA is performed locally instead of globally, on parts of the object deemed inaccurate, using a method based on radius-vector function. Experimental results show that shape PCA reconstruction, especially local shape PCA reconstruction, improves the segmentation in an ore-size measurement application significantly. Maddy Hui Wang, Hong Zhang 0013 |
ICASSP | 2 |
| 2010 | Selection of Gabor filters for improved texture feature extractionabstractTexture feature has been widely used in object recognition, image content analysis and many others. Among various approaches to texture feature extraction, Gabor filter has emerged as one of the most popular ones. Gabor filter-based feature extractor is in fact a Gabor filter bank defined by its parameters including frequencies, orientations and smooth parameters of Gaussian envelope. In the literature, different parameter settings have been suggested, and filter banks created by these parameter settings work well in general. From the perspective of pattern classification, however, filter banks thus designed may not be ideal. In the present study, we propose a new approach to Gabor filter bank design, by incorporating feature selection, i.e. filter selection, into the design process. The merits of incorporating filter selection in filter bank design are twofold. Firstly, filter selection produces a compact Gabor filter bank and hence reduces computational complexity of texture feature extraction. Secondly, Gabor filter bank thus designed produces low-dimensional feature representation with improved sample-to-feature ratio, and this in turn leads to improved performance of texture classification. Experiment results on benchmark datasets and a real application have demonstrated the effectiveness of the proposed method. Kezhi Mao, Hong Zhang 0013, Tianyou Chai |
ICIP | 3 |
| 2010 | Adaptive shape prior in graph cut segmentationabstractIn this paper, we propose a novel method to adaptively apply shape prior in graph cut segmentation. By incorporating shape priors in an adaptive way, we introduce a robust way to harness shape prior in graph cut segmentation. Since traditional graph cut approaches with shape prior may fail in cases where parameters for shape prior term are not set appropriately, incorporation of shape priors adaptively within this framework mitigates these problems. To address this issue, we propose to adaptively apply shape prior based on a shape probability map, defined to reflect the need of shape prior at each location of an image. We show that the proposed method can be easily applied to existing algorithms of graph cut segmentation with shape prior, such as level set based shape prior method, and star shape prior graph cut. We validate our method in various types of images corrupted by significant noise and intensity inhomogeneities. Convincing results are obtained. Maddy Hui Wang, Hong Zhang 0013 |
ICIP | 2 |
| 2010 | Keyframe detection for appearance-based visual SLAMabstractThis paper is concerned with the problem of keyframe detection in appearance-based visual SLAM. Appearance SLAM models a robot's environment topologically by a graph whose nodes represent strategically interesting places that have been visited by the robot and whose arcs represent spatial connectivity between these places. Specifically, we discuss and compare various methods for identifying the next location that is sufficiently different visually from the previously visited location or node in the map graph in order to decide whether a new node should be created. We survey existing techniques of keyframe detection in image retrieval and video analysis. Using experimental results obtained from visual SLAM datasets, we conclude that the feature matching method offers the best performance among five representative methods in terms of accurately measuring the amount of appearance change between robot's views and thus can serve as a simple and effective metric for detecting keyframes. This study fills an important but missing step in the current appearance SLAM research. Hong Zhang 0013, Dan Yang 0001 |
IROS | 1 |
| 2009 | Object Detection with Multiple Motion Models
Zhijie Wang 0003, Hong Zhang 0013 |
ACCV (3) | 2 |
| 2009 | Optimum kernel function design from scale space features for object detectionabstractScale-space representation of an image is a significant way to generate features for classification. However, for a specific classification task, the entire scale-space may not be useful; only a part of it is typically effective. Toward this end, we design a data dependent classification kernel function, which is a weighted mixture of kernels defined on individual scales. In order to choose the optimum weights in the mixture kernel function (MKF), we propose an optimization criterion that leads to the minimization of Raleigh quotient in the positive orthant. This optimization is in general a difficult, non-convex, quadratically constrained quadratic programming. Utilizing a property of ratio of functions, we reduce the aforementioned optimization into a novel binary search, which is essentially a series of quadratic programming. As an application we choose a significant detection problem in oil sands mining called large lump detection from videos. Employing support vector classifier with our MKF yields encouraging results on these difficult-to-process images and compares favorably against the kernel alignment method as well as Fisher criterion adopted in. Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013 |
ICIP | 3 |
| 2009 | Solidity based local threshold for oil sand image segmentationabstractA novel local threshold algorithm for images with poor illumination and complex texture surface is presented in this paper. This algorithm improves segmentation quality by selecting local thresholds according to object level information incorporating prior knowledge, specifically the solidity features. Local thresholds are searched by maximizing the probability of solidity, and fragments with lower segmentation quality are filtered by the stability of solidity. Since thresholding results are produced with object level information, our algorithm is robust in dealing with images of poor quality. Experiments on oil sand images show the proposed algorithm has superior performance to existing local threshold approaches in terms of segmentation quality. Jichuan Shi, Hong Zhang 0013, Nilanjan Ray |
ICIP | 2 |
| 2009 | Tracking of multiple interacting objects using a novel prediction modelabstractTracking multiple interacting objects is an interesting and difficult task in computer vision. Two common problems in this field are a single object with multiple tracks and a single track with multiple objects. Most of the existing algorithms address the first problem but not the second one. In this paper, to solve the second problem we propose a new algorithm with a novel prediction model, which exploits the idea of penalizing outliers in statistics. The experiments show that our proposed algorithm is more robust than the existing algorithms in tackling both the aforementioned problems. Zhijie Wang 0003, Hong Zhang 0013, Nilanjan Ray |
ICIP | 2 |
| 2009 | Development of novel robots with modular methodologyabstractModules have been widely used in the development of re-configurable robots and snake-like robots. Modular methodology can also be applied in design of other robots. To build robots flexibly and quickly with low costs, we have developed two basic joint modules and several functional modules including grippers, suckers and wheels/feet as end-effectors. In this paper, we introduce the development of these modules, and present several novel robots built using them. Specifically, we show how to use them to set up a manipulator, a 6-DoF biped walking robot, a wheeled mobile robot, a biped tree-climbing robot, and a biped wall-climbing robot. It has been shown that a few modules can easily spawn a variety of novel robots with modular methodology. Yisheng Guan, Hong Zhang 0013, Xuefeng Zhou |
IROS | 4 |
| 2009 | Performance evaluation of visual SLAM using several feature extractorsabstractVisual simultaneous localization and mapping (SLAM) implementations must use feature extraction to reduce the dimensionality of image input, yet no comparison of feature extractors exists in the context of visual SLAM. This paper presents both a method for comparison of visual SLAM performance using several different feature extractors and the first experimental study using this method. Possible evaluation metrics are discussed and consistency testing and accumulated uncertainty are chosen to measure performance. Three feature extractors commonly used for visual SLAM are examined: the Harris corner detector, the Kanade-Lucas-Tomasi tracker, and the scale-invariant feature transform. All three are found to perform similarly in an indoor test environment, close to or within the limits of measurement. A modest scale change is handled without difficulty. We conclude that feature extractor choice is not significant in terms of visual SLAM performance and other criteria may be used to make the selection. Jonathan Klippenstein, Hong Zhang 0013 |
IROS | 2 |
| 2009 | An evaluation metric for image segmentation of multiple objects
Mark Polak, Hong Zhang 0013, Ming Hong Pi |
Image Vis. Comput. | 2 |
| 2009 | Ore image segmentation by learning image and shape features
Dipti Prasad Mukherjee, Yury Potapovich, Ilya Levner, Hong Zhang 0013 |
Pattern Recognit. Lett. | 4 |
| 2009 | Snake Validation: A PCA-Based Outlier Detection MethodabstractWe utilize outlier detection by principal component analysis (PCA) as an effective step to automate snakes/active contours for object detection. The principle of our approach is straightforward: we allow snakes to evolve on a given image and classify them into desired object and non-object classes. To perform the classification, an annular image band around a snake is formed. The annular band is considered as a pattern image for PCA. Extensive experiments have been carried out on oil-sand and leukocyte images and the performance of the proposed method has been compared with two other automatic initialization and two gradient-based outlier detection techniques. Results show that the proposed algorithm improves the performance of automatic initialization techniques and validates snakes more accurately than other outlier detection methods, even when considerable object localization error is present. Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013 |
IEEE Signal Process. Lett. | 3 |
| 2008 | Computing oil sand particle size distribution by snake-PCA algorithmabstractAn important measure in various stages of oil sand mining is particle size distribution (PSD) of oil sand particles. Currently PSD is found by time consuming manual inspection. An effective automation of PSD computation can play a significant role in improving the mining process. Toward this goal we propose an algorithm (snake-PCA) to detect oil sands from conveyor belt images, which pose considerable challenges to automated analysis. The novelty in snake-PCA is as follows. First, snake-PCA evolves a number of snakes based on a novel variation of gradient vector flow requiring only a point as initialization. Oil sand is then detected by applying a threshold on PCA reconstruction error of a novel pattern image formed on each evolved snake. We show the discriminative property of the proposed pattern image here. Also, our detection experiments with snake-PCA produce a PSD matching well with a manually found PSD. Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013 |
ICASSP | 3 |
| 2008 | Supervised image segmentation via ground truth decompositionabstractThis paper proposes a data driven image segmentation algorithm, based on decomposing the target output (ground truth). Classical pixel labeling methods utilize machine learning algorithms that induce a mapping from pixel features to individual pixel labels. In contrast we propose to first extract features from both images and labels. Subsequently we induce a mapping from pixel features to label features and synthesize the final output by combining the newly derived label components. We demonstrate the effectiveness of the proposed approach by applying log-Gabor filters to both input and ground truth images of mineral ore. Subsequently we train perceptrons and regression trees to produce individual output components that are combined in frequency space to create the final segmentation. Experimental results show significant improvements over contextual pixel labeling and over ensemble methods. Ilya Levner, Russell Greiner, Hong Zhang 0013 |
ICIP | 3 |
| 2008 | Graph-cut optimization of the ratio of functions and its application to image segmentationabstractOptimizing the ratio of two functions of binary variables is a common task in many image analysis applications. In general, such a ratio is not amenable to graph-cut based optimization. In this paper, we show that if the numerator and the denominator of a ratio are individually graph-representable functions, then their ratio can be optimized via graph-cut based technique. As an example of such a ratio function we choose Yezzi et al.’s energy function [2], minimization of which produces a binary labeling of an image. Through examples, we illustrate the advantage of working with graph-cut-based optimization for the aforementioned ratio in finding a global solution as opposed to the local solutions found by level set methods proposed in [2]. Maddy Hui Wang, Nilanjan Ray, Hong Zhang 0013 |
ICIP | 3 |
| 2008 | Workspace of 3-D multifingered manipulationabstractIn this paper, we propose a numerical approach to generate the workspace of a multifingered robotic hand manipulating an object in the 3-D case. Based on feasibility analysis of grasps, the proposed approach uses a numerical optimization technique to first compute discretely the boundary of the possible motion of the grasped object, and then the limits of rotation about various axes at a specified feasible position of the object. Finally the boundaries of the linear motion and rotation are visualized in 3-D coordinates separately. This approach provides an effective solution to the challenging problem of workspace analysis of 3-D multifingered manipulation. Yisheng Guan, Hong Zhang 0013, Zhangjie Guan |
IROS | 2 |
| 2008 | Consensus-based task sequencing in decentralized multiple-robot systems using local communicationabstractBehavior-based controllers for complex missions often are more easily designed by decomposing the mission into a series of smaller subtasks. When applying this technique to a multiple-robot system, the entire system should focus its work on one subtask at a time to prevent interference from robots working on conflicting subtasks simultaneously. We present a behavior-based approach to the task-sequencing problem for decentralized multiple-robot systems requiring only local inter-robot communication. Consensus is used to prevent premature task sequencing. Our algorithm is compared to globally-communicative and communication-free strategies. The proposed task-sequencing behavior is found to be efficient with respect to time and energy requirements, and could be applied to many existing decentralized multiple-robot systems. Chris A. C. Parker, Hong Zhang 0013 |
IROS | 2 |
| 2007 | A Practical Implementation of Random Peer-to-Peer Communication for a Multiple-Robot SystemabstractIn this paper, we present a physical implementation of random peer-to-peer (RP2P) communication for use in a multiple-robot system and analyze its performance. Traditionally, multiple-robot systems have either broadcast all of their inter-robot communication or have avoided explicit communication altogether. RP2P communication, on the other hand, allows efficient system-level communication while retaining the error-correction capabilities of peer-to-peer connections. We demonstrate that RP2P communication can be implemented with off-the-shelf components. MRS as large as ten robots are investigated and it is demonstrated that message rates as high as 50 messages/second are easily achievable using TCP connections and 802.11B wireless network interfaces. Chris A. C. Parker, Hong Zhang 0013 |
ICRA | 2 |
| 2007 | A UPF-UKF Framework For SLAMabstractIn this paper we propose a SLAM framework which is based on an algorithm that combines an unscented particle filter (UPF) and unscented Kalman filters (UKFs). A UPF is used to estimate robot's poses and the UKFs are used to represent landmark positions. UPF can estimate robot poses more consistently and accurately than generic particle filters (PFs), especially when models are highly non-linear or noises are not Gaussian. UKF can update landmarks more accurately compared to popular EKF's when highly non-linear observation models are used. In addition, our algorithm avoids the calculation of the Jacobian for both motion model and the observation model, which could be extremely difficult for high order systems. The calculation cost of a UPF is on the same order of magnitude as a particle filter (PF), which uses Kalman filters to generate proposal distributions, and the calculation cost of a UKF is equivalent to an EKF. As a result, our SLAM framework is more accurate than other popular SLAM frameworks while its efficiency is maintained. Simulation results are shown to validate the performance goals. Xiang Wang 0008, Hong Zhang 0013 |
ICRA | 2 |
| 2007 | Real-time detection of steam in video images
Ricardo José Ferrari, Hong Zhang 0013, C. Ronald Kube |
Pattern Recognit. | 2 |
| 2007 | Classification-Driven Watershed SegmentationabstractThis paper presents a novel approach for creation of topographical function and object markers used within watershed segmentation. Typically, marker-driven watershed segmentation extracts seeds indicating the presence of objects or background at specific image locations. The marker locations are then set to be regional minima within the topological surface (typically, the gradient of the original input image), and the watershed algorithm is applied. In contrast, our approach uses two classifiers, one trained to produce markers, the other trained to produce object boundaries. As a result of using machine-learned pixel classification, the proposed algorithm is directly applicable to both single channel and multichannel image data. Additionally, rather than flooding the gradient image, we use the inverted probability map produced by the second aforementioned classifier as input to the watershed algorithm. Experimental results demonstrate the superior performance of the classification-driven watershed segmentation algorithm for the tasks of 1) image-based granulometry and 2) remote sensing. Ilya Levner, Hong Zhang 0013 |
IEEE Trans. Image Process. | 2 |
| 2006 | Heuristic Search for Coordinating Robot Agents in Adversarial DomainsabstractThis paper presents a search-based, real-time adaptive solution to the multi-robot coordination problem in adversarial environments. By decomposing the global coordination task into a set of local search problems, efficient and effective solutions to subproblems are found and combined into a global coordination strategy. In turn, each local search entails the use of a heuristic evaluation function together with state space pruning to make the search tractable and scalable. Experimental results, using RoboCup as an example domain, demonstrate the effectiveness of the proposed framework on several simplified RoboCup scenarios Ilya Levner, Alex Kovarsky, Hong Zhang 0013 |
ICRA | 3 |
| 2006 | Color Classification using Adaptive Dichromatic ModelabstractColor-based vision applications face the challenge that colors are variant to illumination. In this paper we present a color classification algorithm that is adaptive to continuous variable lighting. Motivated by the dichromatic color reflectance model, we use a Gaussian mixture model (GMM) of two components to model the distribution of a color class in the YUV color space. The GMM is derived from the classified color pixels using the standard expectation-maximization (EM) algorithm, and the color model is iteratively updated over time. The novel contribution of this work is the theoretical analysis supported by experiments - that a GMM of two components is an accurate and complete representation of the color distribution of a dichromatic surface Xiaohu Lu, Hong Zhang 0013 |
ICRA | 2 |
| 2006 | Bearing-only Landmark Initialization by using SUF with Undistorted SIFT FeaturesabstractIn this paper, we present a delayed algorithm for landmark initialization based on a scaled unscented filter. The algorithm efficiently gives well-conditioned feature locations close to the particle filter based methods even for a very high dimensional nonlinear observation model. Comparing with EKF based methods, our method has the same computational cost while the calculation of Jacobian matrix can be avoided. Experimental results are showed to prove the accuracy and efficiency of the algorithm Xiang Wang 0008, Hong Zhang 0013 |
ICRA | 2 |
| 2006 | Combinatorial Optimization of Sensing for Rule-Based Planar Distributed AssemblyabstractWe describe a model for planar distributed assembly, in which agents move randomly and independently on a two-dimensional grid, joining square blocks together to form a desired target structure. The agents have limited capabilities, including local sensing and rule-based reactive control only, and operate without centralized coordination. We define the spatiotemporal constraints necessary for the ordered assembly of a structure and give a procedure for encoding these constraints in a rule set, such that production of the desired structure is guaranteed. Our main contribution is a stochastic optimization algorithm which is able to significantly reduce the number of environmental features that an agent must recognize to build a structure. Experiments show that our optimization algorithm outperforms existing techniques Jonathan Kelly, Hong Zhang 0013 |
IROS | 2 |
| 2006 | An Analysis of Random Peer-to-Peer Communication for System-Level Coordination in Decentralized Multiple-Robot SystemsabstractInter-robot communication is essential if general purpose intelligent decentralized multiple-robot systems are to become a reality. Traditionally, explicit communication amongst the robots of a MRS has been broadcasted or eliminated altogether. The limits of these two approaches prevent the widespread deployment of useful decentralized MRS. In this paper, we present and analyze, both analytically and empirically, a communication protocol that we call random peer-to-peer or RP2P. Using RP2P, a decentralized MRS can share state in logarithmic time with respect to system population. Additionally, the load placed upon the individual robots by RP2P is independent of system population size. Potential applications of random peer-to-peer communication also are provided, including state sharing, decision-making and task allocation. Chris A. C. Parker, Hong Zhang 0013 |
IROS | 2 |
| 2006 | Good Image Features for Bearing-only SLAMabstractIn this paper, we propose an algorithm for extracting and selecting SIFT (scale-invariant feature transform) visual features for bearing-only SLAM in indoor environments. The algorithm is based on analyzing the stability of the matching ratio of the SIFT features at different scales, and it is capable of extracting SIFT features that can be matched reliably and, at the same time, lead to accurate landmark initialization. In addition, the algorithm is an order of magnitude more efficient than the original SIFT algorithm and is therefore appropriate for the real-time nature of SLAM. As well, the algorithm can determine the quality of the visual features without any delay, and this eliminates the need for a matching or tracking procedure, as is often necessary in other feature extraction algorithms. Results from several experiments verify the performance of the proposed algorithm Xiang Wang 0008, Hong Zhang 0013 |
IROS | 2 |
| 2006 | A Fast and Effective Model for Wavelet Subband Histograms and Its Application in Texture Image RetrievalabstractThis paper presents a novel, effective, and efficient characterization of wavelet subbands by bit-plane extractions. Each bit plane is associated with a probability that represents the frequency of 1-bit occurrence, and the concatenation of all the bit-plane probabilities forms our new image signature. Such a signature can be extracted directly from the code-block code-stream, rather than from the de-quantized wavelet coefficients, making our method particularly adaptable for image retrieval in the compression domain such as JPEG2000 format images. Our signatures have smaller storage requirement and lower computational complexity, and yet, experimental results on texture image retrieval show that our proposed signatures are much more cost effective to current state-of-the-art methods including the generalized Gaussian density signatures and histogram signatures. Ming Hong Pi, Chong-Sze Tong, Siu-Kai Choy, Hong Zhang 0013 |
IEEE Trans. Image Process. | 4 |
| 2005 | Measurement of fine particle size with wavelet signatureabstractThe fine particles of mineral ore are a mixture of granules of various sizes, which collectively look like a texture. It is difficult to estimate the size distribution of such fine particles using conventional image segmentation techniques. In this paper, we propose an image-retrieval-like approach to estimate the size distribution of the fine particles. We introduce a new wavelet signature and a fitting algorithm for wavelet coefficients based on the theory of fragmentation. Ming Hong Pi, Hong Zhang 0013 |
ICIP (3) | 2 |
| 2005 | A rectangular partition algorithm for planar self-assemblyabstractCollective construction is concerned with the study of how multiple robots construct a structure without centralized control. In this paper, we proposed an algorithm to generate labels and transition rules for robots to producing an arbitrary two dimensional structure with only local sensing. The algorithm proceeds by first partitioning the desired structure into a set of rectangles, generating rules and labels for each rectangle, assigning labels to tiles to be placed, generating rules to connect the rectangles, and finally replacing redundant labels and rules in a post-processing step. This algorithm guarantees convergence to the desired structure, produces fewer labels than those algorithms found in the literature, and has the potential to handle structures impossible with a single seed. Hong Zhang 0013 |
IROS | 2 |
| 2005 | Broker: an interprocess communication solution for multi-robot systemsabstractWe describe in this paper a novel implementation of the interprocess communication (IPC) technology, called Broker, in support of the development and the operation of a complex robot system. We view each robot system as a collection of processes that need to exchange information, e.g. motion commands and sensory data, in a flexible and convenient fashion, without affecting each other's operations in case of a process's scheduled termination or unexpected failure. We argue that the IPC technology provides an ideal framework for this purpose, and we carefully make our design decisions about its implementation based on the needs of robotics applications. Broker is programming language, operating system, and hardware platform independent and has served us well in a RoboCup project and collective robotics experiments, in both simulation and real-world environments. Matthew McNaughton, Sean Verret, Andrzej Zadorozny, Hong Zhang 0013 |
IROS | 4 |
| 2005 | Distance based communication in the surveillance task in a multi-robot systemabstractWe investigate communication effects in the surveillance task and compare a communication vs. no-communication condition in a simulated environment using a modified version of an existing tracking algorithm. The modification not only allowed the collective to achieve a steady state of 100% performance in most cases but also reduced both the influence of target-to-robot ratios and the size of the work radius as well as the performance variance. Communication solved an interference effect, rather than create one. The results are discussed from a perspective that focuses on messages that move the collective as a whole from an initial state to a solution state. Such messages provide information about the current task state, rather than information about the state of individual members or the location of attractors. Rabih Neouchi, Hong Zhang 0013, René Elio |
IROS | 2 |
| 2005 | Active versus passive expression of preference in the control of multiple-robot decision-makingabstractJust like their solitary counterparts, multiple-robot systems must be able to make decisions in response to their environment. However, with a multiple-robot system, one must take care to ensure that the individual robots that compose a system make their decisions in concert with each other. We desire decisions to be made at the system level. In this paper, we investigate four different mechanisms to allow individual robots within a system to express their preference for a particular solution to a system-level problem. All four mechanisms consistently produced unanimous decisions, but had varying ability to produce unanimous decisions of good quality. An approach that we refer to as passive expression of preference performed the best, but had to be tuned to the particular problem being solved. A mechanism that we refer to as active expression of preference exhibited very good performance and required no problem specific tuning, which makes it more universally applicable to the multiple-robot decision-making problem. Chris A. C. Parker, Hong Zhang 0013 |
IROS | 2 |
| 2005 | Modified GMM background modeling and optical flow for detection of moving objectsabstractSegmentation of moving objects in image sequences is a fundamental step in many computer vision applications such as mineral processing industry and automated visual surveillance. In this paper, we introduce a novel approach to detect moving objects in a noisy background. Our approach combines a modified adaptive Gaussian mixture model (GMM) for background subtraction and optical flow methods supported by temporal differencing in order to achieve robust and accurate extraction of the shapes of moving objects. The algorithm works well for image sequences having many moving objects with different sizes as demonstrated by experimental results on real image sequences. Dongxiang Zhou, Hong Zhang 0013 |
SMC | 2 |
| 2005 | A multistage adaptive thresholding method
Feixiang Yan, Hong Zhang 0013, C. Ronald Kube |
Pattern Recognit. Lett. | 2 |
| 2004 | Characterization of acuity laser range finderabstractThis paper presents an experimental evaluation of the performance of a laser range finder, AccuRange 4000 by Acuity Research, a popular distance sensor in robotics research and industrial automation. Our study determines the performance of the ranging device under various operating conditions, including lighting, temperature, and surface color, and orientation. Based on the empirical data, a calibration model is developed to allow one to interpret the range measurements accurately. Xiujuan Luo, Hong Zhang 0013 |
ICARCV | 2 |
| 2004 | Coaching a Robot CollectiveabstractThis paper is concerned with a study in which the advantages of telerobotics and supervisory control are combined with those of a multi-robot system, in the context of robot soccer. We propose our solutions to two important issues in such systems, namely, the level of abstraction at which the commands to a robot collective are issued, and the selection of the group within the collective to direct the commands. The design philosophy as well as its implementation are described. We show through experiments that, statistically, a supervised robot soccer team out-performs an unsupervised team, with the help of coaching advices. The proposed approach has applications beyond robotic soccer. Marc Perron, Hong Zhang 0013 |
ICRA | 2 |
| 2004 | Biologically inspired decision making for collective robotic systemsabstractPractical collective robotic systems likely will be confronted with problems which have more than one unique solution. When deciding on which of a set of candidate solutions to a problem to pursue, a collective system should ensure that its members reach a unanimous decision regarding which solution to implement so that the system itself does not split apart with different members pursuing different solutions. If such a split were to occur, much of the collective system's functionality could be lost. In this paper, we present a unique approach to collective decision making that is based on an algorithm employed by a particular species of ant when it chooses a new nest site. We expand the ants' algorithm into a general purpose decision making scheme and apply it to the collective relocation problem. A detailed study of the performance of our decision making algorithm was carried out in simulation using the collective relocation task as a test bed. Consistent system performance was observed across three robot populations. It was found that one particular system variable, the decision quorum threshold played a large role in determining the system's behaviour and that system behaviour was maximized when this variable was set to 50% of the system's population. Chris A. C. Parker, Hong Zhang 0013 |
IROS | 2 |
| 2004 | Collective sorting with local communicationabstractThis paper describes an experimental study of the problem of collective sorting, in which multiple robots under simple reactive control rearrange initially randomly distributed objects of different classes into separate clusters so that each cluster contains objects of only one class. We specifically focus on the roles of three key system parameters - namely, population size, local communication among robots, and sensing range - on the performance of the collective sorting task of objects of two classes. Our experiments show that an increase in communication range has a similar positive effect on task performance as an increase in sensing range, especially when the sensing range is limited; however, this effect tends to be attenuated by the increase in the size of the robot collective. Sean Verret, Hong Zhang 0013, Max Q.-H. Meng |
IROS | 2 |
| 2004 | Latest research in computer vision - a special issue on VI 2002
Hong Zhang 0013, Dmitry O. Gorodnichy |
Image Vis. Comput. | 1 |
| 2003 | Image-based localization with depth-enhanced image mapabstractIn this paper, we present an image-based robot incremental localization algorithm which uses a panoramic image-based map enhanced with depth from a laser range finder. The image-based map (model) contains both intensity information as well as sparse 3D geometric features. By assuming motion continuity, a robot can use the depth information in the image-model to project the relevant 3D model features, specifically vertical lines, of the environment to its camera coordinate frame. To determine its location, the robot first acquires an intensity image and then matches the 2D geometric features in the image with the projected model features. The first contribution of this research is that we avoid the difficult problem of full 3D reconstruction from images by employing a range sensor registered with respect to the intensity image sensor; secondly, we provide an algorithm that performs incremental robot localization using only 2D images. Experimental results in indoor map building and localization demonstrate the feasibility of our approach and evaluate the performance of the algorithm. Dana Cobzas, Hong Zhang 0013, Martin Jägersand |
ICRA | 2 |
| 2003 | Feasibility analysis of 2D graspsabstractIn this paper, we develop a novel approach to evaluating the feasibility of a planar grasp of an arbitrary polygonal object by a multifingered hand. Our definition of a feasible grasp takes into account both kinematic constraints and force constraints, and the algorithm is capable of handling power grasps involving finger tips as well as inner links of a dextrous hand. The proposed approach is based on first expressing the kinematic and force constraints in terms of equalities and inequalities in variables of the object and hand configurations, and then formulating and solving grasp feasibility as a constrained global optimization problem. An example is provided to illustrate the algorithm. Yisheng Guan, Hong Zhang 0013 |
IROS | 2 |
| 2003 | Workspace of 2D multifingered manipulationabstractThe knowledge of the workspace of a multi-fingered hand is very important in planning a dexterous manipulation task. In this paper, we propose a novel approach to compute and visualize the workspace of a multifingered robotic hand manipulating an object in the planar case. Based on optimization models, our approach is numerical, in which the kinematic feasibility of a grasp at a given position is determined first, and then the range of rotation at this position is computed. The algorithm and its effectiveness are illustrated by examples. It is possible to extend our approach to the spatial case. Yisheng Guan, Hong Zhang 0013 |
IROS | 2 |
| 2003 | Blind bulldozing: multiple robot nest constructionabstractIn this paper, we present a collective, or swarm construction algorithm to control robotic bulldozers in the creation of a work site. Predictions about robotic missions to the planet Mars have described such site preparation as essential to the success of later mission objectives, such as the construction of solar arrays, etc. This algorithm was based on a behaviour observed in a particular species of ant called "blind bulldozing". We developed a mathematical model of blind bulldozing using a unique approach based on Markov chains. Robot bulldozers were developed and used to test the algorithms in our laboratory. The team of robots was found to be successful at clearing an open area out of field of rocks. Our robots' behaviour also agreed with the predictions of our model. This work is significant because it demonstrated the viability of blind bulldozing and represents the first time, to our knowledge, that a multiple robot system has carried out a form of the general construction task outside of simulation. Chris A. C. Parker, Hong Zhang 0013, C. Ronald Kube |
IROS | 2 |
| 2003 | Kinematic feasibility analysis of 3-D multifingered graspsabstractPlanning of a dextrous manipulation task for a multifingered hand requires the feasibility of all the grasps involved throughout the manipulation process. In this paper, we address the problem of determining whether a desired grasp of a polyhedral object is kinematically feasible. In our study, we define a grasp in terms of a system of contact pairs between the topological features of the hand and the object, and formulate the grasp feasibility analysis as a set of equality and inequality constraints in the variables of the hand and object configurations. The feasibility of a grasp then becomes equivalent to the simultaneous satisfaction of all the constraints. This allows us to cast the feasibility analysis conveniently as a constrained nonlinear optimization problem and solve it numerically with commercially available software. The effectiveness of our approach is illustrated with an example of grasping a cuboid using a three-fingered robotic hand. Yisheng Guan, Hong Zhang 0013 |
IEEE Trans. Robotics Autom. | 2 |
| 2002 | A Comparative Analysis of Geometric and Image-Based Volumetric and Intensity Data Registration AlgorithmsabstractWe present and contrast four methods for registering 3D range data to 2D images. Two are calibration techniques that recover the rigid transformation between the sensor poses, based on point or line correspondences. The two others recover a direct, image based mapping between the data sets. The accuracy of each method is experimentally evaluated on test patterns and objects. We found that the point based calibration method is the best approach to recover a global registration between the two sensors, while an image-based method performed best when registering local regions. Dana Cobzas, Hong Zhang 0013, Martin Jägersand |
ICRA | 2 |
| 2002 | An indexing scheme for efficient data-driven verification of 3D pose hypotheses
Michael Boshra, Hong Zhang 0013 |
Image Vis. Comput. | 2 |
| 2002 | Hybrid resistive tactile sensingabstractThe paper describes a novel design of a robot tactile sensor, called hybrid tactile sensor, based on analog resistive technology. The design employs a set of parallel analog resistive sensing strips, rather than a two-dimensional (2D) grid of sensing elements, to sample the sensor surface. It is therefore analog in one dimension and discrete in the other. As a result, the number of samples that need to be obtained is significantly reduced, as is the number of connectors in the sensor circuitry. In addition, the sensor design scales easily to a large sensing surface area, and offers flexibility in sensor geometry. However, only the contact location and shape can be sensed, but not the contact force. The paper provides the theoretical model of the sensor, and presents simulation and experimental results. Hong Zhang 0013, Eric So |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2001 | An Integrated Robotic Hand/Simulator System for Tele-manipulation via the InternetabstractTo enhance the dexterity of a tele-manipulation system and to provide a more efficient and intuitive user interface, we have developed an integrated system for remote dextrous manipulation using a multifingered robotic hand through the Internet. This system consists of a three-fingered hand for real-time execution and a graphic simulator for manipulation simulation and command input. We describe the main issues in the development of this system, including system architecture, functions, integration and user interface. We also provide some experiments of remote manipulation of typical objects with this system. Our work verifies the feasibility of such an integrated system, and provides a new approach to the potential application of a multifingered robotic hand. Yisheng Guan, Teresa Ho, Hong Zhang 0013 |
ICRA | 3 |
| 2001 | Kinematic Feasibility Analysis of 3D GraspsabstractWe present a solution to the problem of determining if it is possible for a given dexterous hand to grasp a polyhedral object at a desired topological configuration, subject to kinematic constraints. We refer to this problem as kinematic feasibility analysis. In our study, we define a desired grasp in terms of a set of contact pairs between the topological features of the hand and the object, and formulate a general algorithm that makes use of constrained optimization with nonlinear constraints to determine the feasibility. We first derive the conditions in order for the hand to make the desired contact and avoid undesirable collision. These conditions are expressed in terms of equalities and inequalities in the configuration variables of the hand and object, which can then easily be transformed into a constrained nonlinear optimization problem. Numerical examples are provided to illustrate the solution. Yisheng Guan, Hong Zhang 0013 |
ICRA | 2 |
| 2001 | Hand Pose Recovery with a Single Video CameraabstractWe present a system to determine human hand pose from a single monochrome image in real-time. We formulate the hand pose problem as determining the position and orientation of the palm as well as the joint angles of the fingers. The 3D pose of the palm and each joint angle are computed simultaneously using feature points on the human hand. We exploit an existing fast object pose algorithm developed by DeMenthon and Davis (1995) together with a heuristic minimization of a 1D function. We make use of a special glove to remove noise caused by the deformation of the hand and to help find correspondence between 3D object points and their images. We demonstrate our system by using a mock-up cardboard hand with one finger. Experimental results with a real human hand wearing the special glove are presented. Kihwan Kwon, Hong Zhang 0013, Fadi Dornaika |
ICRA | 2 |
| 2001 | Cylindrical panoramic image-based model for robot localizationabstractMobile robots must be able to learn, use and maintain models of their environments. Two fundamental and interrelated problems in mobile robotics are mapping and localization. We present here a new type of image-based map formed by panoramic models enriched with depth and 3D planarity information. We take an advantage of the scene geometry implicitly contained in the model to localize a mobile robot that is moving in the same environment. The novelty of this model compared to the existing image-based maps is that the motion of the robot is not restricted to a predefined path or locations close to the original images. Experimental results demonstrate this approach. Dana Cobzas, Hong Zhang 0013 |
IROS | 2 |
| 2000 | Kinematic Graspability of a 2D Multifingered HandabstractWe describe a solution to the problem of determining the kinematic feasibility of a dextrous hand to grasp a given object at a desired configuration in the two dimensional space. We refer to this problem as kinematic graspability. The kinematic configuration of a grasp is defined in terms of a set of contact pairs between the topological features of the hand and the object. We derive the sufficient conditions in order for the hand to make the desired contact and avoid collision. These conditions are then formulated as a constrained nonlinear global optimization problem, whose solution yields a definitive answer to kinematic graspability for a grasp configuration. Numerical examples are provided to illustrate the method. Yisheng Guan, Hong Zhang 0013 |
ICRA | 2 |
| 2000 | Hybrid video using motion estimationabstractOne of the major problems in low-bandwidth telerobotics applications is determining what visual data is sufficient to allow the operator to successfully perform a task. When the available bandwidth is extremely low, we must severely restrict the data being transmitted. For very-low-bandwidth applications, displaying certain types of features, such as edges in indoor navigation applications, allows for successful completion of the operator's task. As the available bandwidth increases, the operator can be presented with more information than just the edge features. Using simple overlay techniques to overlay a low-frame-rate video stream on a high-frame-rate edge feature stream provides the operator with more information about the environment. Registering the video data with the moving edge data poses a significant problem. In this paper, we propose to use simple motion estimation techniques based on the edge image stream to estimate the motion of the edge images and use this data to register the slower frame-rate video stream to the edge feature stream. Using these simple estimation techniques allows us to perform the video stream registration quickly as the edge feature data is being presented to the operator. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
SMC | 3 |
| 2000 | Localizing a polyhedral object in a robot hand by integrating visual and tactile data
Michael Boshra, Hong Zhang 0013 |
Pattern Recognit. | 2 |
| 2000 | Control of contact via tactile sensingabstractWe present our approach to using tactile sensing in the feedback control of robot contact tasks. A general framework, called tactile servo, is first introduced. We then address the critical issue of how to model the state of contact in terms that are both sufficient for defining general contacts and conducive to bridging the gap between a robot task description and the information observable by a tactile sensor. We subsequently examine techniques for deriving tactile sensor models that are required for computing the state of contact from the sensor output. Two basic methods for mapping a tactile sensor image to the contact state variables are introduced: one based on explicit inverse modeling, and the other on numerical tactile Jacobian. In both cases, moments of the tactile sensor image are used as the features that capture the necessary information about contact. The theoretical development is supported by extensive experiments, which include edge tracking, object shape construction and object manipulation. Hong Zhang 0013, Ning Nicholas Chen |
IEEE Trans. Robotics Autom. | 1 |
| 1999 | Panoramic Video with Predictive Windows for Telepresence ApplicationsabstractWe describe the application of a predictive Kalman filter to the display of panoramic images. We discuss integrating a panoramic imaging system with prediction of viewing direction to create an effective telepresence system over low bandwidth links. Panoramic imaging using a reflective mirror surface offers an alternative to pan-tilt systems for obtaining a 360 degree field of view. Selecting a small window within a panoramic image allows a meaningful part of an image from a remote site to be seen at a higher refresh rate. Because of the delay in transmitting an image from a remote site, it is necessary to have additional image information available locally. This information can be used to simulate continuously flowing pictures with reduced apparent delay. Continuity in image viewing is achieved by predicting the next viewpoint of an operator and preemptively transmitting parts of an image. Experimental results are given to evaluate the proposed telepresence system. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
ICRA | 3 |
| 1999 | A Constraint-Satisfaction Approach for 3-D Object Recognition by Integrating 2-D and 3-D Data,
Michael Boshra, Hong Zhang 0013 |
Comput. Vis. Image Underst. | 2 |
| 1999 | Accommodating uncertainty in pixel-based verification of 3-D object hypotheses
Michael Boshra, Hong Zhang 0013 |
Pattern Recognit. Lett. | 2 |
| 1998 | Predictive Windows for Delay Compensation in Telepresence ApplicationsabstractPredictive Kalman filters can be used to predict positions of a mouse when it is operated under some basic assumptions. This prediction can be used to estimate what portions of a larger image an operator wants to view. This paper discusses theory and experimentation being done at the University of Alberta using predictive Kalman filters to provide predictive windows for low bandwidth telepresence applications. We compare several state models used in prediction with each other and also with having no prediction, both with numerical measures, and on human subjects. We show that the constant velocity model provides the best prediction results. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
ICRA | 3 |
| 1997 | Touch-driven robot control using a tactile JacobianabstractAn experimental study on touch-driven robot control, or tactile-servo, is presented in this paper. The tactile servo scheme uses a tactile feature Jacobian matrix to relate the differential change in tactile feature space to that in the robot task space. Tactile Jacobian matrices are constructed through a finite element (FE) model of the tactile sensor. Detailed implementation of the control scheme and experimental results are presented. Hong Zhang 0013, Raymond E. Rink |
ICRA | 2 |
| 1996 | An efficient pixel-based technique for visual verification of 3-D object hypothesesabstractThe technique proceeds in three steps: firstly, the visible-edge image of the hypothesized model object is constructed; secondly, this image is superimposed on the scene edge image; finally, corresponding pixels in the two images are compared to gather votes for the validity of the hypothesis. Uncertainty in estimating the locations of both scene and model edge pixels is handled by dilating the scene edge image. An analytical method is presented for determining the extent of dilation, assuming some error bound on the object pose. The proposed technique has several important advantages over the common feature-based techniques. Firstly, it is much simpler to implement. Secondly, the time complexity of verifying a hypothesis is made independent of the scene complexity (number and types of scene features). Thirdly, it is insensitive to imperfections of the feature extraction process. Finally the technique is easy to implement on the existing parallel vision/graphics hardware; thus it is suitable for real-time applications. Michael Boshra, Hong Zhang 0013 |
ICRA | 2 |
| 1996 | Local object shape from tactile sensingabstractStudies on contact problems such as shape recovery from tactile sensing involve complex geometrics and kinematics. In the first part of the paper we introduce a convenient matrix expression of 3-D surfaces, which can be easily applied to manipulate contact constraints and kinematics equations using homogeneous transformation. In the second part of the paper we apply the notation introduced in the first part and derive an active tactile sensing strategy which consists of a shape sensing algorithm and a touch-based robot control scheme. Raymond E. Rink, Hong Zhang 0013 |
ICRA | 3 |
| 1996 | The use of perceptual cues in multi-robot box-pushingabstractIn this paper we present an approach to controlling transitions in multirobot tasks which have been modelled as a linear series of steps. A box-pushing task is described as a sequence of sub-tasks with a separate controller designed for each step using finite state automata theory. Perceptual cues are formed by concatenating binary variables which represent locally sensed stimuli into boolean vectors used to specify transitions between sub-task steps. The approach is designed for a redundant set of homogeneous mobile robots equipped with simple sensors and stimulus-response behaviours. A set of perceptual cues used in box-pushing are designed and tested on 10 physical mobile robots. It is argued that perceptual cues and finite state automata offers a new approach to environment-specific task modelling in collective robotics. C. Ronald Kube, Hong Zhang 0013 |
ICRA | 2 |
| 1996 | Sensitivity analysis and experiments of curvature estimation based on rolling contactabstractCurvature is a useful feature for characterizing arbitrary smooth 3-D surfaces. This paper describes our study of detecting object curvature using tactile sensing. Using the constraints on the relative motion between two bodies in contact, the curvature of an object can be obtained by inducing a relative differential motion and sensing the change in the location of contact. To understand the limitation of the practical application of this approach, sensitivity of the curvature estimate to errors in sensory data is presented. The validity of the theoretical development is examined experimentally using an optical wave-guide based tactile sensor and one finger of a three-finger dextrous hand. A good agreement between the theoretical analysis and the experimental results has been obtained. Hong Zhang 0013, Hitoshi Maekawa, Kazuo Tanie |
ICRA | 1 |
| 1996 | Dextrous manipulation planning by grasp transformationabstractA novel framework for dextrous manipulation planning based on transformations between canonical grasp configurations is proposed in this paper. The proposed approach is inspired by the observation that human hand is capable of a large number of manipulative tasks with a few finger postures or grasps. It is postulated that the key to the control of manipulation tasks, therefore, lies in the combination of these grasps. The proposed framework first identifies important or canonical grasps, defined in terms of contact pairs between topological features of the hand and the object, and then obtains the possible transitions between the grasps. With the result expressed in the form of a grasp transformation graph, general manipulation tasks can then be planned by searching the graph for a path connecting the initial and final configurations of the tasks. The proposed approach is demonstrated by experiments using a three-finger hand performing object manipulation tasks. Hong Zhang 0013, Kazuo Tanie, Hitoshi Maekawa |
ICRA | 1 |
| 1995 | A constraint-satisfaction approach for 3D vision/touch-based object recognitionabstractWe present a technique for recognizing polyhedral objects by integrating visual and tactile data. The problem is formulated as a constraint-satisfaction problem (CSP) to provide a unified framework for integrating different types of sensory data. To make use of the scene perceptual structures early in the recognition process, we enforce local consistency of the CSP. The process of local-consistency enforcing (LCE) reduces the correspondence uncertainty between scene and model features, which can lead to significant reductions in the computational load on subsequent recognition modules. LCE can also eliminate many erroneous model objects efficiently, without explicitly generating or verifying any object/pose hypotheses. Michael Boshra, Hong Zhang 0013 |
IROS (2) | 2 |
| 1995 | Efficient edge detection from tactile dataabstractA real-time tactile image processing algorithm for edge contact is presented in this paper. Based on basic elasticity results, closed-form solutions for calculating contact force and local contact geometries (i.e., location and orientation of the line of contact) from the first three moments of a tactile image are derived. Computational complexity of the proposed algorithm and those of the previous approaches are compared, and passive tactile sensing experiments are performed. It is shown that the proposed algorithm has the advantage of persevering force information and is more consistent and computationally efficient. Raymond E. Rink, Hong Zhang 0013 |
IROS (3) | 3 |
| 1995 | Edge tracking using tactile servoabstractA novel method to perform edge tracking using tactile sensors is presented in this paper. Using the tactile servo scheme, a robot manipulator is driven only by real-time tactile feedback from the array tactile sensors mounted directly on the robot end-effector. Compared with previous approaches, the control scheme presented in this paper is consistent and more efficient. Real-time edge tracking experiments are conducted using an experimental system consisting of a PUMA 260, a single rigid finger and a planar array tactile sensor. Experimental results show satisfactory control speed and accuracy for both straight and curved edge tracking. An example of active tactile sensing of a unknown object using edge tracking is also demonstrated. Hong Zhang 0013, Raymond E. Rink |
IROS (2) | 2 |
| 1995 | Two-dimensional optimal sensor placementabstractA method for determining the optimal two-dimensional spatial placement of multiple sensors participating in a robot perception task is introduced in this paper. This work is motivated by the fact that sensor data fusion is an effective means of reducing uncertainties in sensor observations, and that the combined uncertainty varies with the relative placement of the sensors with respect to each other. The problem of optimal sensor placement is formulated and a solution is presented in two dimensional space. The algebraic structure of the combined sensor uncertainty with respect to the placement of sensors is studied. A necessary condition for optimal placement is derived and this necessary condition is used to obtain an efficient closed-form solution for the global optimal placement. Numerical examples are provided to illustrate the effectiveness and efficiency of the solution.> Hong Zhang 0013 |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1994 | Crack Detection Using Contact SensingabstractIn this paper, we describe a robotic system employing contact sensing to detect the presence of cracks in surfaces. Automated detection of cracks in surfaces has many practical applications. Tasks such as inspection of underground pipes carrying fluids and the inspection of ship hulls are examples of jobs that are difficult for human operators. Robots equipped with vision are ill-suited for such tasks because of low visibility and fluid disturbance. Contact sensing is shown to provide a viable alternative for surface inspection tasks in situations where images cannot be obtained. The mathematical modeling of the system, some simulations and experimental results are provided.> Raju Patil, Anup Basu, Hong Zhang 0013 |
ICRA | 3 |
| 1994 | Use of visual and tactile data for generation of 3-D object hypothesesabstractMost existing 3-D object recognition/localization systems rely on a single type of sensory data, although several sensors may be available in a robot task to provide information about the objects to be recognized. In this paper, the authors present a technique to localize polyhedral objects by integrating visual and tactile data. It is assumed that visual data is provided by a monocular visual sensor, while tactile data by a planar-array tactile sensor in contact with the object to be localized. The authors focus on using tactile data in the hypothesis generation phase to reduce the requirements of visual features for localization to a V-junction only. The main concept of this technique is to compute a set of partial pose hypotheses off-line by utilizing tactile data, and then complement these partial hypotheses on-line using visual data. The technique presented is tested using simulated and real data.> Michael Boshra, Hong Zhang 0013 |
IROS | 2 |
| 1994 | Stagnation recovery behaviours for collective roboticsabstractAccomplishing useful tasks with a collection of decentralized mobile robots will require control methods that deal effectively with a number of unique problems that impede the system's progress. Reactive control architectures can easily cause the problems of stagnation and cyclic behaviour, both characterized by a lack of progress in achieving a task. In this paper the authors present one possible solution to stagnation recovery, motivated from the study of group transport in ants and demonstrate its use in a box-pushing task. By using stagnation recovery behaviours, which are triggered by a lack of progress in the task-achieving activity of the system, the collective system can monitor its own advancement in a decentralized manner. A set of such behaviours are progressively ordered using timeouts, with each set designed for a specific recovery strategy. The stagnation recovery behaviours have been tested in simulation with the results to be mapped onto a set of ten autonomous robots presently under construction.> C. Ronald Kube, Hong Zhang 0013 |
IROS | 2 |
| 1993 | Controlling collective tasks with an ALNabstractIn this paper, the authors explore the idea of using an adaptive logic network (ALN) for behavior arbitration. Their approach is to define the collective task, to be performed by multiple robots, as a group behavior. The group behavior is a set of behaviors, each of which specifies a single step in the collective task. An environmental cue can be used to control the transition between behaviors, thus allowing the progress of the collective task to self-govern its execution. The authors simplify the behavior arbitration in the collective task by training an ALN implemented using simple combinational logic. They provide a description of the Collective Robotic Intelligence Project (CRIP) including the simulation results and the multi-robot system on which these results will be deployed. C. Ronald Kube, Hong Zhang 0013, Xiaohuan Wang |
IROS | 2 |
| 1993 | Efficient evaluation of the feasibility of robot displacement trajectoriesabstractA robot motion trajectory is feasible if it lies entirely in the robot's workspace. It is often desirable to be able to evaluate efficiently if a motion trajectory is feasible so as to terminate infeasible trajectories before their execution begins. A method that allows one to determine the feasibility of displacement trajectories of wrist-partitioned robot manipulators is described. The workspace description of a robot manipulator in general using workspace analysis based on polynomial discriminants is obtained, and the simplicity of the workspace description by examples is shown. A set of algebraic equations is derived that allows one to evaluate the feasibility of displacement trajectories efficiently without using inverse kinematics.> Hong Zhang 0013 |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1992 | Optimal sensor placementabstractThe author defines the problem of optimal sensor placement in two-dimensional space to address the issue of choosing a configuration of multiple sensors to minimize the uncertainty in the consensus estimate obtained by combining the sensor measurements. The algebraic structure of sensor errors with respect to sensor placement is studied, and a necessary condition for the optimal placement is derived. The solution to the problem is obtained by comparing sensor configurations that satisfy the necessary condition. The method does not assume specific sensor modality and is applicable to sensors of any type that are capable of providing position information.> Hong Zhang 0013 |
ICRA | 1 |
| 1992 | Motion Planning With Acceleration ConstraintabstractIn this paper we examine the problem of finding a local collision free path between two points in a two dimensional space. We consider a non-h.olonomic robot with constraints on normal acceleration. Solution, of this problem is important in order to prevent a robot moving on wheels from slipping (or skidding) while turning. The exact solution (known so far) to the problem of reachability is exponential. We analyze th,e problem of finding an approximate path locally with an a priori probability (or approximation measure). Our algorithm is based on a variable size discretization, which is computed locally depending on the given constraints, the size of the local free space, and the closeness of the approximation. Implemen,tation results demonstrating the validity of the method are presented. Anup Basu, Goksin Bakir, Hong Zhang 0013 |
IROS | 3 |
| 1991 | Feasibility analysis of displacement trajectories for robot manipulators with a spherical wristabstractThe author describes a method by which to determine the feasibility of displacement trajectories of a class of robot manipulators, namely, those with a wrist of intersecting axes. The author obtains the workspace description of a robot manipulator in general using workspace analysis based on polynomal discriminants, and shows the simplicity of the workspace description of the PUMA 560 in particular. Taking advantage of that simplicity, he then shows that there exist convenient algebraic equations that require minimum computation and enable the feasibility of displacement trajectories to be evaluated in some cases without using inverse kinematics.> Hong Zhang 0013 |
ICRA | 1 |
| 1991 | A parallel inverse kinematics solution for robot manipulators based on multiprocessing and linear extrapolationabstractThe authors present a method to compute inverse kinematics in parallel for robots with a closed-form solution, and distinguish a pipelined solution from a parallel solution. Although both increase the system throughput, only the parallel solution reduces the computational latency. The computational task of computing inverse kinematics is partitioned with one subtask per joint, and all subtasks are computed in parallel. The intrinsic dependency among subtasks is removed by linear extrapolation through the gradient of the inverse kinematic functions and joint velocity information. The simplicity of the solution makes it easily applicable to any robot manipulator with a closed-form solution. Examples are used to illustrate the effectiveness and efficiency of the algorithm. Implementation of the algorithm on a multiprocessor system is described.> Hong Zhang 0013, Richard P. Paul |
IEEE Trans. Robotics Autom. | 1 |
| 1990 | A parallel inverse kinematics solution for robot manipulators based on multiprocessing and linear extrapolationabstractA method of computing inverse kinematics in parallel for robots with a closed-form solution is presented. The computational task of computing each inverse kinematics solution is partitioned with one subtask per joint, and all subtasks are computed concurrently. The intrinsic dependency among subtasks is removed by linear extrapolation through the gradient of inverse kinematic functions and joint velocity information. The high degree of concurrency and naturally balanced concurrent subtasks of the system significantly reduce the latency of the inverse kinematics evaluation. Compared with a serial solution, the algorithm results in a reduction of the time of execution by a factor proportional to the number of joints when implemented on a multiprocessor system. Its simplicity makes it easily applicable to any robot manipulators with closed-form solutions. Examples are used to illustrate the effectiveness and the efficiency of the algorithm. Implementation of the algorithm on a multiprocessor system is also discussed.> Hong Zhang 0013, Richard P. Paul |
ICRA | 1 |
| 1989 | Kinematic stability of robot manipulators under force controlabstractThe author investigates the kinematic instability problem associated with some well-known force control approaches. His intent is to establish qualitative evaluations of the stability of hybrid control and stiffness control. He first constructs a state-space model and then demonstrates the kinematic stability or instability of the two force control formulations by making use of a necessary and sufficient condition for a multivariate linear system. The author also discusses issues such as critical damping in stable force control approaches and stabilizing measures for unstable force control methods.> Hong Zhang 0013 |
ICRA | 1 |
| 1988 | Non-kinematic errors in robot manipulatorsabstractThe ability of a robot manipulator to position its end-effector accurately depends not only on kinematic modeling of the robot but also on the dynamic execution of the transformation between end-effector coordinates and joint coordinates. The authors examine errors associated with various approaches of motion trajectory generation and evaluate them statistically. Two approaches to controlling errors are considered: VAL, which has been developed by them, and resolved motion rate control (RMRC), which is due to D.E. Whitney (1986).> Hong Zhang 0013, Richard P. Paul |
ICRA | 1 |
| 1988 | A parallel solution to robot inverse kinematicsabstractIntroduces an algorithm by which the inverse kinematics of a robot manipulator with closed-form solution can be computed in parallel to reduce the computational complexity roughly by a factor of n, the number of joints of the manipulator. The authors study the errors introduced by the algorithm by constructing statistical models of the errors. They evaluate the algorithm on a PUMA 260 manipulator, which has six degrees of freedom with six revolute joints and a reach of approximately 0.5 m.> Hong Zhang 0013, Richard P. Paul |
ICRA | 1 |
| 1986 | Design of a robot force/motion serverabstractThe design and implementation of a distributed robot manipulator controller based on a parallel architecture is introduced in this paper. We consider the robot manipulator as a robot force and motion server (RFMS) to the robot system to execute force and motion commands issued by the robot coordinator. In the server, computation is distributed to a number of processors to achieve a system capable of executing computation expensive tasks. The use of "c" programming language for both the system programming and the user interface provides a convenient environment to control such a system. Richard P. Paul, Hong Zhang 0013 |
ICRA | 2 |
| 1985 | Hybrid control of robot manipulatorsabstractA robot manipulator must be able to satisfy any consistent position and force specifications in a well-defined Cartesian coordinate frame. In this paper, we present a method to achieve such a hybrid control of position and force. While the desired compliance is specified in Cartesian space, the control is accomplished in joint space, thus effectively reducing the computation load. Results show that accurate position and force control is obtained. Hong Zhang 0013, Richard P. Paul |
ICRA | 1 |