VLDB 2026 Research / reviewers in the wild / expert
Dezhen Song
dblp:16/3547
· DBLP profile ↗
101ranked-venue papers
17as first author
25since 2021 · last 2026
0000-0002-2944-5754ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 10 first-author · 15 since 2021Systems, architecture and hardware · 59 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4 · 4 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSR-Net: Crack-Aware Style Recomposition Network for Airport Runways
Nansha Li, Haifeng Li 0008, Zhongcheng Gui, Dezhen Song |
ICIC (21) | 5 |
| 2025 | A Full-Optical Pretouch Dual-Modal and Dual-Mechanism (PDM2) Sensor for Robotic GraspingabstractWe report a new full-optical pretouch dual-modal and dual-mechanism (PDM2) sensor based on an air-coupled fiber-tip surface micromachined optical ultrasound transducer (SMOUT). Compared to ring-shaped piezoelectric acoustic receivers in previous PDM2sensors, the acoustic signal received by the new fiber-tip SMOUT is readout optically, which is naturally resistant to surrounding electromagnetic interference (EMI) and makes the complex grounding and shielding unnecessary. In addition, the new fiber-tip SMOUT receiver has a much smaller size, which makes it possible to further miniaturize the sensor package into a more compact structure. For verification, a prototype of the full-optical PDM2sensor has been designed, fabricated, and characterized. The experimental results show that even with the much smaller acoustic receiver, the new sensor can still achieve ranging and material/structure sensing performances comparable with the previous ones. Therefore, the new full optical PDM2sensor design is promising in providing a practical and miniaturized solution for ranging and material/structure sensing to assist robotic grasping of unknown objects. Zhiyu Yan, Fengzhi Guo, Shuangliang Li, Dezhen Song |
ICRA | 5 |
| 2025 | Heterogeneous Sensor Fusion and Active Perception for Transparent Object Reconstruction with a PDM2 Sensor and a CameraabstractTransparent household objects present a challenge for domestic service robots, since neither regular cameras nor RGB-D cameras can provide accurate points for shape reconstruction. The new type of pretouch dual-modality distance and material sensor (PDM2) can provide reliable and accurate depth readings, but it is a point sensor and scanning the object exclusively with the sensor is too inefficient. Hence, we present a sensor fusion approach by combining a regular camera with the PDM2sensor. The approach is based on a data fusion algorithm for shape reconstruction and an active perception algorithm for scan planning for the PDM2sensor. The data fusion algorithm is a distributed Gaussian process (GP)-based shape reconstruction method that allows for incremental local update to reduce computational time. The active perception algorithm is an optimization-based approach by increasing the information gain (IG) and prioritizing the boundary points under a preset travel distance constraint. We have implemented and tested the algorithms with six different transparent household items. The results show satisfactory shape reconstruction results in all test cases with an average increase in intersection over union (IoU) from 0.73 to 0.96. Fengzhi Guo, Shuangyu Xie, Di Wang 0020, Dezhen Song |
ICRA | 6 |
| 2025 | Energy Efficient Planning for Repetitive Heterogeneous Tasks in Precision AgricultureabstractRobotic weed removal in precision agriculture introduces a repetitive heterogeneous task planning (RHTP) challenge for a mobile manipulator. RHTP has two unique characteristics: 1) an observe-first-and-manipulate-later (OFML) temporal constraint that forces a unique ordering of two different tasks for each target and 2) energy savings from efficient task collocation to minimize unnecessary movements. RHTP can be framed as a stochastic renewal process. According to the Renewal Reward Theorem, the expected energy usage per task cycle is the long-run average. Traditional task and motion planning focuses on feasibility rather than optimality due to the unknown object and obstacle position prior to execution. However, the known target/obstacle distribution in precision agriculture allows minimizing the expected energy usage. For each instance in this renewal process, we first compute task space partition, a novel data structure that computes all possibilities of task multiplexing and its probabilities with robot reachability. Then we propose a region-based setcoverage problem to formulate the RHTP as a mixed-integer nonlinear programming. We have implemented and solved RHTP using Branch-and-Bound solver. Compared to a baseline in simulations based on real field data, the results suggest a significant improvement in path length, number of robot stops, overall energy usage, and number of replans. Shuangyu Xie, Kenneth Y. Goldberg, Dezhen Song |
ICRA | 3 |
| 2025 | Ground Penetrating Radar-Assisted Multimodal Robot Odometry Using Subsurface Feature MatrixabstractLocalization of robots using subsurface features observed by ground-penetrating radar (GPR) enhances and adds robustness to common sensor modalities, as subsurface features are less affected by weather, seasons, and surface changes. We introduce an innovative multimodal odometry approach using inputs from GPR, an inertial measurement unit (IMU), and a wheel encoder. To efficiently address GPR signal noise, we introduce an advanced feature representation called the subsurface feature matrix (SFM). The SFM leverages frequency domain data and identifies peaks within radar scans. Additionally, we propose a novel feature matching method that estimates GPR displacement by aligning SFMs. The integrations from these three input sources are consolidated using a factor graph approach to achieve multimodal robot odometry. Our method has been developed and evaluated with the CMU-GPR public dataset, demonstrating improvements in accuracy and robustness with real-time performance in robotic odometry tasks. Haifeng Li 0008, Jiajun Guo, Xuanxin Fan, Huaichao Wang, Kairat Koshekov, Dezhen Song |
ICTAI | 7 |
| 2025 | Towards Safe Imitation Learning via Potential Field-Guided Flow MatchingabstractDeep generative models, particularly diffusion and flow matching models, have recently shown remarkable potential in learning complex policies through imitation learning. However, the safety of generated motions remains overlooked, particularly in complex environments with inherent obstacles. In this work, we address this critical gap by proposing Potential Field-Guided Flow Matching Policy (PF2MP), a novel approach that simultaneously learns task policies and extracts obstacle-related information, represented as a potential field, from the same set of successful demonstrations. During inference, PF2MP modulates the flow matching vector field via the learned potential field, enabling safe motion generation. By leveraging these complementary fields, our approach achieves improved safety without compromising task success across diverse environments, such as navigation tasks and robotic manipulation scenarios. We evaluate PF2MP in both simulation and real-world settings, demonstrating its effectiveness in task space and joint space control. Experimental results demonstrate that PF2MP enhances safety, achieving a significant reduction of collisions compared to baseline policies. This work paves the way for safer motion generation in unstructured and obstacle-rich environments. Anqing Duan, Zezhou Sun, Leonel Rozo, Noémie Jaquier, Dezhen Song, Yoshihiko Nakamura |
IROS | 6 |
| 2025 | Simulating Automotive Radar with Lidar and Camera InputsabstractLow-cost millimeter automotive radar has received more and more attention due to its ability to handle adverse weather and lighting conditions in autonomous driving. However, the lack of quality datasets hinders research and development. We report a new method that is able to simulate 4D millimeter wave radar signals including pitch, yaw, range, and Doppler velocity along with radar signal strength (RSS) using camera image, light detection and ranging (lidar) point cloud, and ego-velocity. The method is based on two new neural networks: 1) DIS-Net, which estimates the spatial distribution and number of radar signals, and 2) RSS-Net, which predicts the RSS of the signal based on appearance and geometric information. We have implemented and tested our method using open datasets from 3 different models of commercial automotive radar. The experimental results show that our method can successfully generate high-fidelity radar signals. Moreover, we have trained a popular object detection neural network with data augmented by our synthesized radar. The network outperforms the counterpart trained only on raw radar data, a promising result to facilitate future radar-based research and development. Peili Song, Dezhen Song, Enfan Lan, Jingtai Liu |
IROS | 2 |
| 2024 | AR-Classroom Usability: Implications for UX Research on AR-Enabled Educational Technologies for 3D Matrix Algebra LearningabstractThis research full paper describes the augmented reality (AR) application AR-Classroom that combines a physical and virtual environment to teach 3D geometric rotations in an engaging and simplified manner. The AR-Classroom contains a virtual workshop where users can perform rotations by manipulating the application's X, Y, and Z-axis sliders to rotate a virtual LEGO model and a physical workshop where users perform rotations using a physical LEGO model. Guided by previous findings and an iterative approach to usability, the present usability study focused on assessing the usability of the AR-Classroom in its most recent version using a new physical LEGO model (i.e., airplane) and reflecting on how discoverability and usability can be assessed using different types of user experience measures and qualitative analysis. Participants were 22 undergraduate students who completed a pre-test with demographic information, watched a video on geometric transformations, and were randomly assigned to interact with either workshop. While interacting with the AR app, participants were instructed to provide feedback and a single ease-of-use question score. Participants then completed a post-test with two measures of usability. Descriptive statistics of the UX measures were explored, and a thematic analysis was conducted to identify and code themes in human-computer interaction. Findings suggest that the AR-Classroom has reached satisfactory usability, and users can navigate the app's features effectively. Discussion includes insight into which aspects of the app need improvement, how to promote self-directed support in the app, the development of future app efficacy experiments, and how to evaluate the usability of AR technology for learning. Samantha D. Aguilar, Chengyuan Qian, Uttamasha Monjoree, Heather Burte, Jeffrey Liew, Francis K. H. Quek, Philip Yasskin, Dezhen Song, Wei Yan 0006 |
FIE | 8 |
| 2024 | AR-Classroom: Integrating Conversational Artificial Intelligence with Augmented Reality Technology for Learning Spatial Transformations and Their Matrix RepresentationabstractThis research full paper describes the AR-Classroom application that utilizes augmented reality (AR) and physical and virtual manipulatives to enable undergraduate students to build intuition about the relation between spatial transformations and their mathematical representations. To further build on the app's usability and functionality, additional features are being prototyped to continue improving the user-app interaction with the AR-Classroom. Some of the challenges the students faced when using AR-Classroom were recalling basic matrix operations without geometric context, basic trigonometric functions and their applications in the two-dimensional space, loss of AR registration for not understanding the AR environment, and User Interface (UI) issues. To address these issues, a conversational Artificial Intelligence (AI)-based multi-sensory and interactive assistance has been added to the AR-Classroom. Integrating sophisticated language processing and response generation of AI with immersive three-dimensional capabilities of AR can create a more engaging learning experience than the previous versions of the app. This integration focuses on creating a symbiosis between AR and AI. It creates an elevated user experience by offering real-time, personalized assistance to students dealing with issues related to understanding mathematical concepts and functionalities of the app. A qualitative exploratory usability study was done to assess the user's interaction with the AI implemented in the AR-Classroom, aiming to explore the AI's ability to guide students in using AR technology and aid in introductory matrix algebra learning, to effectively serve the students' learning. Based on the thematic analysis of the user experiment we found four main themes related to users' perceptions of AR-Classroom AI features usability: (1) AI chatbot ease-of-use, (2) Need for answer elaboration from AI, (3) Desire for visual information, and (4) Increased understanding of the content area. The scores of ease of use indicate AI's ability to guide complex tasks in an AR environment using AI features with less concern for the cognitive load. The overall result suggests the need for further investigation on incorporating AI-guided visual cues in an AR environment. Uttamasha Monjoree, Samantha D. Aguilar, Chengyuan Qian, Carl Van Huyck, Shu-Hao Yeh, Preston Tranbarger, Luke Duane-Tessier, Leo Solitare-Renaldo, Heather Burte, Philip Yasskin, Jeffrey Liew, Dezhen Song, Francis Quek, Wei Yan 0006 |
FIE | 12 |
| 2024 | Subsurface Feature-based Ground Robot/Vehicle Localization Using a Ground Penetrating RadarabstractRobot localization using subsurface features captured by Ground-Penetrating Radar (GPR) complements and improves robustness over existing common sensor modalities, as subsurface features are less sensitive to weather, season and surface scene changes. Here, we propose a novel subsurface feature-based localization method that uses only GPR measurements with a known subsurface map. An efficient feature descriptor, the dominant energy curve (DEC), is designed to identify different locations in cluttered conditions. Specifically, image processing techniques that involve background segmentation, energy point detection, and energy curve refinement are designed to extract DEC features from a 2D radargram. With DECs features obtained, a metric subsurface feature map is constructed. Finally, we perform robot localization by feature matching under a particle swarm optimization framework. We have implemented our method and tested it with the public CMU-GPR dataset. The results show that our algorithm improves accuracy and robustness with real-time performance for robot localization tasks. Specifically, the mean localization error is 0.50 m for all cases. Haifeng Li 0008, Jiajun Guo, Dezhen Song |
ICRA | 3 |
| 2024 | Coupled Active Perception and Manipulation Planning for a Mobile Manipulator in Precision Agriculture ApplicationsabstractA mobile manipulator often finds itself in an application where it needs to take a close-up view before performing a manipulation task. Named this as a coupled active perception and manipulation (CAPM) problem, we model the uncertainty in the perception process and devise a key state/task planning algorithm that considers reachability conditions jointly established from perception and manipulation task constraints. By minimizing expected energy usage in body key state planning while satisfying task constraints, our algorithm is able to find an energy-efficient trajectory with less body repositioning motion while ensuring the success of the task. We have implemented the algorithm and tested it in both simulation and physical experiments. The results have confirmed that our algorithm has a lower energy consumption compared to a two-stage decoupled approach, while still maintaining a success rate of 100% for the task. Shuangyu Xie, Chengsong Hu, Di Wang 0020, Joe Johnson, Muthukumar Bagavathiannan, Dezhen Song |
ICRA | 6 |
| 2024 | Toward Precise Robotic Weed Flaming Using a Mobile Manipulator with a BlowtorchabstractRobotic weed flaming is a new and environmentally friendly approach to weed removal in the agricultural field. Using a mobile manipulator equipped with a blowtorch, we design a new system and algorithm to enable effective weed flaming, which requires robotic manipulation with a soft and deformable end effector, as the thermal coverage of the flame is affected by dynamic or unknown environmental factors such as gravity, wind, atmospheric pressure, fuel tank pressure, and pose of the nozzle. System development includes overall design, hardware integration, and software pipeline. To enable precise weed removal, the greatest challenge is to detect and predict dynamic flame coverage in real time before motion planning, which is quite different from a conventional rigid gripper in grasping or a spray gun in painting. Based on the images from two onboard infrared cameras and the pose information of the blowtorch nozzle on a mobile manipulator, we propose a new dynamic flame coverage model. The flame model uses a center-arc curve with a Gaussian cross-section model to describe the flame coverage in real time. The experiments have demonstrated the working system and shown that our model and algorithm can achieve a mean average precision (mAP) of more than 76% in the reprojected images during online prediction. Di Wang 0020, Chengsong Hu, Shuangyu Xie, Joe Johnson, Hojun Ji, Yingtao Jiang, Muthukumar Bagavathiannan, Dezhen Song |
IROS | 8 |
| 2024 | Road Boundary Estimation Using Sparse Automotive Radar InputsabstractLow-cost millimeter wavelength automotive radar can work effectively under low visibility or low reflection conditions caused by lighting, weather, pollution, or object surface properties when a camera or a lidar may fail. It can serve as a fallback solution to improve safety in autonomous driving. However, after filtering, radar signals tend to be sparse and noisy which poses new challenges in scene understanding. This paper presents a new approach to detecting road boundaries based on sparse radar signals. We model the roadway using a homogeneous model and derive its conditional predictive model under known radar motion. Using this predictive model and modeling radar points using a Dirichlet Process Mixture Model, we employ Mean Field Variational Inference (MFVI) to derive an unconditional road boundary model distribution. To generate initial candidate solutions for the MFVI, we develop a custom Random Sample and Consensus (RANSAC) variant to propose unseen model instances as candidate road boundaries. For each radar point cloud we alternate the MFVI and RANSAC proposal steps until convergence to generate the best estimate of all candidate models. We select the candidate model with the minimum lateral distance to the radar on each side as the estimates of the left and right boundaries. We have implemented the proposed algorithm and it has shown satisfactory results. More specifically, the mean lane boundary estimation error is not more than 11.0 cm. Aaron Kingery, Dezhen Song |
IROS | 2 |
| 2023 | AR-Classroom: Usability of AR Educational Technology for Learning Rotations Using Three-Dimensional Matrix AlgebraabstractThe AR-Classroom application utilizes augmented reality technology (AR) to make the three-dimensional (3D) rotations underlying matrix algebra visible and interactive. The AR-Classroom has physical and virtual versions, where users can perform rotations using a physical LEGO model or by manipulating the application's x, y, and z axes sliders to rotate a virtual model. Both versions provide 3D matrices, color-coded axes lines, and a green wireframe superimposed onto a LEGO model to represent transformations. To ensure that the AR-Classroom makes learning 3D matrix algebra more engaging and accessible, two usability tests were used to evaluate the discoverability and usability of the app. The benchmark test assessed usability in the AR-Classroom's original format, and recommendations were made to improve the app., such as adding additional instructions on model set-up, restructuring, and updating the instructions, and turning the 'visualization type 'function into a button to make it easier to find. After the improvements, the updated usability test assessed usability again so that the impact of the modifications could be evaluated. Participants followed similar procedures in both the benchmark$(\mathrm{N}=12)$and updated usability$(\mathrm{N}=12)$tests. Participants completed a pre-test assessing their math abilities and confidence, watched a video on geometric transformations, and then were randomly assigned to interact with either the physical or virtual version of the app. While interacting with the app, participants were given tasks to complete while thinking out loud and provided an ease-of-use rating from$1=\text{very}$easy to$7=\text{very}$difficult (i.e., SEQ score). Once done interacting with the app, participants completed a post-test assessing their math abilities and confidence and provided feedback on their overall experience with the app (i.e., SUS). A thematic analysis was conducted after each test to identify and code themes in interaction and compare findings from the benchmark and updated tests. Results indicated that after changes were made to the app, the usability of both versions significantly improved: users were better able to set up the space shuttle model, effectively utilize the in-app instructions, and quickly access all of the app's features. Findings from the updated usability test contribute to enhancing the AR-Classroom app and further its use in higher education classrooms for learning matrix algebra. Samantha D. Aguilar, Heather Burte, Philip Yasskin, Jeffrey Liew, Shu-Hao Yeh, Chengyuan Qian, Dezhen Song, Uttamasha Monjoree, Wei Yan 0006 |
FIE | 7 |
| 2023 | AR-Classroom: Augmented Reality Technology for Learning 3D Spatial Transformations and Their Matrix RepresentationabstractProject AR-Classroom aims to enhance undergraduate students learning spatial transformations and their mathematical representations. Understanding closely allied spatial and mathematical concepts significantly contributes to STEM learning in fields of computer graphics, computer-aided design, computer vision, robotics, and many more. The technology and learning innovations of this research include novel AR features and their implications for learning. In AR-Classroom, a student can hold and manipulate a 3D physical model (a LEGO space shuttle as an example) while simultaneously interacting with AR visualization of 3D rotations. Two usability tests with 24 participants total have been conducted for AR-Classroom leading to promising results and recommendations for improvements. The project contributes to advancing our knowledge in (1) the role of interplay between physical and virtual manipulatives to engage students in embodied learning and (2) the features of AR to make difficult, invisible concepts visible for supporting an intuitive and formal understanding of spatial reasoning and mathematical formulation. Shu-Hao Yeh, Chengyuan Qian, Dezhen Song, Samantha D. Aguilar, Heather Burte, Philip Yasskin, Ziad Ashour, Zohreh Shaghaghian, Uttamasha Monjoree, Wei Yan 0006 |
FIE | 3 |
| 2023 | The Third Generation (G3) Dual-Modal and Dual Sensing Mechanisms (DMDSM) Pretouch Sensor for Robotic GraspingabstractFingertip-mounted pretouch sensors are very useful for robotic grasping. In this paper, we report a new (G3) dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor for near-distance ranging and material sensing, which is based on pulse-echo ultrasound (US) and optoacoustics (OA). Different from previously reported versions, the G3 sensor utilizes a self-focused US/OA transceiver, thereby eliminating the need of a bulky parabolic reflective mirror for focusing the ultrasound and laser beams. The self-focused laser and ultrasound beams can be easily steered by a (flat) scanning mirror which expands from single-point ranging and detection to areal mapping or imaging. To verify the new design, a prototype G3 DMDSM sensor with a scanning mirror is fabricated. The US and OA ranging performances are tested in experiments. Together with the scanning mirror, thin wire targets made of same or different materials at different positions are scanned and imaged. The ranging and imaging results show that the G3 DMDSM sensor can provide new and better pretouch mapping and imaging capabilities for robotic grasping than its predecessors. Shuangliang Li, Di Wang 0020, Fengzhi Guo, Dezhen Song |
ICRA | 5 |
| 2023 | A Pretouch Perception Algorithm for Object Material and Structure Mapping to Assist Grasp and Manipulation Using a DMDSM SensorabstractWe report a new material and structure mapping (MSM) algorithm to assist robotic grasping and manipulation. Building on our new sensor development, the algorithm has four main components: 1) detection of time-of-flight (ToF) durations for the dual modalities of optoacoustic (OA) and pulse-echo ultrasound (US), 2) contour reconstruction by fusing OA and US signals, 3) local noise filtering by checking local consistency of material and structure label (MSL), and 4) medium boundary searching that identifies class boundaries through two-staged clustering and boundary establishment using support vector machine (SVM) hyperplanes. We have implemented our algorithm and tested it with multiple common household items. The experimental results have successfully validated our algorithm design which shows that the average error of contour reconstruction is 0.05 mm and the true positive rate of MSL is over 98%. Fengzhi Guo, Shuangyu Xie, Di Wang 0020, Dezhen Song |
IROS | 6 |
| 2023 | On Perpendicular Curve-Based Task Space Trajectory Tracking Control With Incomplete Orientation ConstraintabstractThe Incomplete Orientation Constraint (IOC), which does not require a controlled motion constrained by all three spatial directions, exists widely in a lot of robotic tasks. However, the IOC remains a challenge for existing methods, due to the nonlinear structure of the rotation group SO(3). Moreover, the IOCs are time varying in the trajectory tracking problem, which makes it more challenging than the set-point control. To address the IOC problems, we define, identify and prove the closed-form solution of the perpendicular curve in SO(3). Based on the proposed perpendicular curve, we develop a new trajectory tracking controller considering the IOC. Compared with existing methods, the proposed method can achieve faster and more accurate tracking results. Moreover, the proposed method can be applied not only to manipulators with redundancy (including both functional and intrinsic redundancy), but also to manipulators that are non-redundant. Furthermore, it is easier to incorporate a secondary optimization objective into consideration for the intrinsic redundant case, which is difficult for existing methods. The proposed method has been implemented on both simulations and experiments. The numerous simulation and experimental results validate the effectiveness and advantages of the proposed method. Note to Practitioners—The trajectory tracking control with IOC is important and can be applied in a wide spectrum of robotic-assisted manufacturing, e.g. arc-welding, engraving, etc.. In this paper, a perpendicular curve-based trajectory tracking method is proposed to consider the IOC automatically. It is no longer necessary to carefully plan a reachable orientation trajectory. Users only need to focus the planning of the tool direction, which is determined by the tasks. In addition, the proposed method can achieve faster and more accurate tracking result. Moreover, it is universal and can be applied to both redundant and non-redundant cases. Gaofeng Li, Shan Xu 0002, Dezhen Song, Fernando Caponetto, Ioannis Sarakoglou, Jingtai Liu, Nikolaos G. Tsagarakis |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2023 | M2FNet: Multimodal Fusion Network for Airport Runway Subsurface Defect Detection Using GPR DataabstractGround penetrating radar (GPR) is widely used for detecting airport runway subsurface defects. The trailing interference in the GPR data disguises subsurface defect responses, which seriously affects the accuracy of subsurface defect detection. To tackle the challenge of subsurface defects detection under trailing interference in real scenarios, a new multi-modal fusion network referred to as M2FNet is proposed. Based on the premise that the trailing signal is highly similar across all adjacent A-scans, the model employs a transformer encoder to extract global features of the signal with long-distance correlation. In contrast, the subsurface defects only show echo characteristics in a few adjacent B-scans. The phase of the trailing and the target signal is opposite, which is easy to be discovered from the Top-scan view. Thus, a hybrid convolutional neural network structure is used to extract local features from GPR images of different views. This dual network structure extremely enhance the representation learning of GPR data. In order to investigate various subsurface defects and trailing interference collected by multiple GPR systems under different conditions, the first large-scale hybrid dataset called ASD-GPR is created. Transfer learning is employed to enhance the model’s ability to detect rare defects by fine-tuning it for real-world situations, differing from synthetic training data scenarios. The results of the experiments reveal that M2FNet outperforms state-of-the-art object detection methods in various real-world scenarios, demonstrating superior performance in detecting subsurface defects. Nansha Li, Renbiao Wu, Haifeng Li 0008, Huaichao Wang, Zhongcheng Gui, Dezhen Song |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Gaussian-Process-Based Control of Underactuated Balance Robots With Guaranteed PerformanceabstractThe control of underactuated balance robots is aimed at performing both the external (actuated) subsystem trajectory tracking and internal (unactuated) subsystem balancing tasks. In this article, we propose a learning-based control design for underactuated balance robots. The key idea integrates a model predictive control method to design the desired internal subsystem trajectory and perform the external subsystem tracking task, while an inverse dynamics controller is used to stabilize the internal subsystem to its desired trajectory. The control design is based on Gaussian process (GP) regression models that are learned from experiments without requiringa prioriknowledge about the robot dynamics or the demonstration of successful stabilization. GP regression models also provide estimates of modeling uncertainties of the robotic systems, and these estimations are used to enhance control robustness to modeling errors. The learning-based control design is analyzed with guaranteed stability and performance. The proposed design is demonstrated by experiments on a Furuta pendulum and an autonomous bikebot. Kuo Chen, Jingang Yi, Dezhen Song |
IEEE Trans. Robotics | 3 |
| 2022 | The Second Generation (G2) Fingertip Sensor for Near-Distance Ranging and Material Sensing in Robotic GraspingabstractTo continuously improve robotic grasping, we are interested in developing a contactless fingertip-mounted sensor for near-distance ranging and material sensing. Previously, we demonstrated a dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor prototype based on pulse-echo ultrasound and optoacoustics. However, the complex system, the bulky and expensive pulser-receiver, and the omni-directionally sensitive microphone block the sensor from practical applications in real robotic fingers. To address these issues, we report the second generation (G2) DMDSM sensor without the pulser-receiver and microphone, which is made possible by redesigning the ultrasound transmitter and receiver to gain much wider acoustic bandwidth. To verify our design, a prototype of the G2 DMDSM sensor has been fabricated and tested. The testing results show that the G2 DMDSM sensor can achieve better ranging and similar material/structure sensing performance, but with much-simplified configuration and operation. The primary results indicate that the G2 DMDSM sensor could provide a promising solution for fingertip pretouch sensing in robotic grasping. Di Wang 0020, Dezhen Song |
ICRA | 3 |
| 2021 | Device Design and System Integration of a Two-Axis Water-immersible Micro Scanning Mirror (WIMSM) to Enable Dual-modal Optical and Acoustic Communication and Ranging for Underwater VehiclesabstractTo address the communication and ranging challenges caused by underwater environment, we design dual modal devices for autonomous underwater vehicles (AUVs). The dual-modal design builds upon a co-axial ultrasonic and green laser beams which leverage different signal diverging patterns and different responses in the underwater environment by each modality to achieve robust adaptability. Here we report our recent progress in improving scanning and aiming capabilities for dual-modal beam steering. The core part is our Two-Axis Water-immersible Micro Scanning Mirror (WIMSM). We improve hinge design of WIMSM for larger scanning range. We incorporate high speed Hall effect sensor-based pose feedback channel to enable closed-loop scanning and aiming control. We design ultrasonic-assisted laser handshaking method to help AUVs to acquire optical underwater communication. We have prototyped our devices and tested them in a water tank. The initial results are promising. Xiaoyu Duan, Di Wang 0020, Dezhen Song |
ICRA | 3 |
| 2021 | Fingertip Pulse-Echo Ultrasound and Optoacoustic Dual-Modal and Dual Sensing Mechanisms Near-Distance Sensor for Ranging and Material Sensing in Robotic GraspingabstractTo improve robotic grasping, we are interested in developing a new non-contact fingertip-mounted sensor for near-distance ranging and material sensing. Here we report new progress in combining direct pulse-echo ultrasound and optoacoustic effects in sensor design to deal with optically and/or acoustically challenging targets (OACTs). Our dual-modal and dual sensing mechanisms (DMDSM) sensor design is enabled by a novel wideband ultrasound transmitter embedded inside a piezoelectric (lead zirconate titanate - PZT) ring transducer. The new DMDSM sensor is capable of differentiating a variety of OACTs. To verify our design, both distance ranging tests and material sensing tests have been conducted. The ranging tests show the sensor can perform both optoacoustic ranging (for light-absorbing materials) and pulse-echo ultrasound ranging (for reflective or transparent materials). For material sensing, the dual-modal spectra from OACTs are collected to compare the new sensor with previous designs. The overall 100% accuracy from the confusion matrices indicates the initial success of our sensor design in differentiating conventional targets as well as the OACTs with the new DMDSM sensor. Di Wang 0020, Dezhen Song |
ICRA | 3 |
| 2021 | Graph-Based Proprioceptive Localization Using a Discrete Heading-Length Feature Sequence Matching ApproachabstractProprioceptive localization refers to a new class of robot egocentric localization methods that do not rely on the perception and recognition of external landmarks. These methods are naturally immune to bad weather, poor lighting conditions, or other extreme environmental conditions that may hinder exteroceptive sensors such as a camera or a laser ranger finder. These methods depend on proprioceptive sensors such as inertial measurement units and/or wheel encoders. Assisted by magnetoreception, the sensors can provide a rudimentary estimation of vehicle trajectory which is used to query a prior known map to obtain location. Named as graph-based proprioceptive localization, we provide a low cost fallback solution for localization under challenging environmental conditions. As a robot/vehicle travels, we extract a sequence of heading-length values for straight segments from the trajectory and match the sequence with a preprocessed heading-length graph (HLG) abstracted from the prior known map to localize the robot under a graph-matching approach. Using the information from HLG, our location alignment and verification module compensates for trajectory drift, wheel slip, or tire inflation level. We have implemented our algorithm and tested it in both simulated and physical experiments. The algorithm runs successfully in finding robot location continuously and achieves localization accurate at the level that the prior map allows (less than 10 m). Hsin-Min Cheng, Dezhen Song |
IEEE Trans. Robotics | 2 |
| 2021 | Encoder-Camera-Ground Penetrating Radar Sensor Fusion: Bimodal Calibration and Subsurface MappingabstractIn this article, we report system and algorithmic developments for a sensing suite comprising a camera and a ground penetrating radar (GPR) with a wheel encoder designed for both surface and subsurface infrastructure inspection, which is a multimodal mapping task. To fuse different sensor modalities properly, we solve a novel GPR-camera calibration problem and a synchronization-challenged sensor fusion problem. First, we design a calibration rig, model the GPR imaging process, introduce a mirror to obtain the joint coverage between the camera and the GPR, and employ the maximum-likelihood estimator to estimate the relative pose between the camera and the GPR with error analysis. Second, we propose a data collection scheme using the customized artificial landmarks to synchronize camera images (temporally evenly spaced) and GPR/encoder data (spatially evenly spaced). We also employ pose graph optimization with location discrepancy as penalty functions to perform data fusion for 3-D reconstruction. We have tested our system in physical experiments. The results show that our system successfully fuses encoder-camera-GPR sensory data and accomplishes a metric 3-D reconstruction. Moreover, our sensor fusion approach reduces the end-to-end distance error from 6.4 to 0.7 cm in a real bridge inspection experiment if comparing to the counterpart that only uses encoder measurements. Chieh Chou, Haifeng Li 0008, Dezhen Song |
IEEE Trans. Robotics | 3 |
| 2020 | Fingertip Non-Contact Optoacoustic Sensor for Near-Distance Ranging and Thickness Differentiation for Robotic Grasping*abstractWe report the feasibility study of a new optoacoustic sensor for both near-distance ranging and material thickness classification for robotic grasping. It is based on the optoacoustic effect where focused laser pulses are used to generate wideband ultrasound signals in the target. With a much smaller optical focal spot, the optoacoustic sensor achieves a lateral resolution of 93 μm, which is six times higher than ultrasound pulse-echo ranging under the same condition. A new multi-mode wideband PZT (lead zirconate titanate) transducer is built to properly receive the wideband optoacoustic signal. The ability to receive both low- and high-frequency components of the optoacoustic signal enhances the material sensing capability, which makes it promising to determine not only material type but also the sub-surface structures. For demonstration, optoacoustic spectra are collected from hard and soft materials with different thickness. A Bag-of-SFA-Symbols (BOSS) classifier is designed to perform primary material and then thickness classification based on the optoacoustic spectra. The accuracy of material / thickness classification reaches ≥ 99% and ≥ 94%, respectively, which shows the feasibility of differentiating solid materials with different thickness by the optoacoustic sensor. Di Wang 0020, Dezhen Song |
IROS | 3 |
| 2020 | Lane Marking Verification for High Definition Map Maintenance Using Crowdsourced ImagesabstractAutonomous vehicles often rely on high-definition (HD) maps to navigate around. However, lane markings (LMs) are not necessarily static objects due to wear & tear from usage and road reconstruction & maintenance. Therefore, the wrong matching between LMs in the HD map and sensor readings may lead to erroneous localization or even cause traffic accidents. It is imperative to keep LMs up-to-date. However, frequently recollecting data with dedicated hardware and specialists to update HD maps is not only cost-prohibitive but also unviable. Here we propose to utilize crowdsourced images from multiple vehicles at different times to help verify LMs for HD map maintenance. We obtain the LM distribution in the image space by considering the camera pose uncertainty in perspective projection. Both LMs in HD map and LMs in the image are treated as observations of LM distributions which allow us to construct posterior conditional distribution (a.k.a Bayesian belief functions) of LMs from either sources. An LM is consistent if belief functions from the map and the image satisfy statistical hypothesis testing. We further extend the Bayesian belief model into a sequential belief update using crowdsourced images. LMs with a higher probability of existence are kept in the HD map whereas those with a lower probability of existence are removed from the HD map. We verify our approach using real data. Experimental results show that our method is capable of verifying and updating LMs in the HD map. Binbin Li 0006, Dezhen Song, Aaron Kingery, Dongfang Zheng, Yiliang Xu, Huiwen Guo |
IROS | 2 |
| 2020 | Model Quality Aware RANSAC: A Robust Camera Motion EstimatorabstractRobust estimation of camera motion under the presence of outlier noisevision. Despite existing efforts that focus on detecting motion and scene degeneracies, the best existing approach that builds on Random Consensus Sampling (RANSAC) still has non-negligible failure rate. Since a single failure can lead to the failure of the entire visual simultaneous localization and mapping, it is important to further improve the robust estimation algorithm. We propose a new robust camera motion estimator (RCME) by incorporating two main changes: a model-sample consistency test at the model instantiation step and an inlier set quality test that verifies model-inlier consistency using differential entropy. We have implemented our RCME algorithm and tested it under many public datasets. The results have shown a consistent reduction in failure rate when comparing to the RANSAC-based Gold Standard approach and two recent variations of RANSAC methods. Shu-Hao Yeh, Dezhen Song |
IROS | 3 |
| 2020 | Toward Automatic Subsurface Pipeline Mapping by Fusing a Ground-Penetrating Radar and a CameraabstractWe propose a novel subsurface pipeline mapping and 3D reconstruction method by fusing ground-penetrating radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with nonperpendicular angles. Furthermore, we fuse visual simultaneous localization and mapping outputs, encoder readings with GPR scans to classify hyperbolas into different pipeline groups. We extensively apply the J-linkage method and maximum likelihood estimation with error analysis to improve algorithm robustness and accuracy. As a result, we optimally estimate the radii and locations of all pipelines. We have implemented our method and tested it in physical experiments with representative pipeline configurations. Two different kinds of 3-m-long pipes are used, with radii being 4.62 and 3.02 cm, respectively. The results show that our method successfully reconstructs all subsurface pipes. Moreover, the average estimation errors for two orientation angles of pipelines are 1.73° and 0.73°, respectively. The average localization error is 4.47 cm. Note to Practitioners-Automatic and accurate underground pipeline mapping technology is very important in civil construction projects. Lack of 3D utility pipeline maps may lead to accidental damage in civil construction and maintenance. Although ground-penetrating radar (GPR-based pipeline mapping methods have been studied for several years, these methods require the perpendicular scanning with respect to the pipe, which is impossible to guarantee in practice since the orientations of pipelines are unknown. Furthermore, these traditional methods can only estimate one pipeline at a time in a survey area and require prior knowledge of pipe diameter. We propose a robotic subsurface pipeline mapping method with a GPR and a camera to handle difficult factors such as multiple pipes, unknown pipeline orientation, and unknown pipeline diameters. Hence, we can perform GPR scanning along any generic linear trajectories. Our method has been tested in physical experiments with representative pipeline configurations. The results are sufficiently accurate, and it proves that our method can be an effective technology to reconstruct the underground pipelines. Haifeng Li 0008, Chieh Chou, Longfei Fan, Binbin Li 0006, Di Wang 0020, Dezhen Song |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2019 | Gaussian Processes Model-Based Control of Underactuated Balance RobotsabstractControl of underactuated balance robot requires external subsystem trajectory tracking and internal unstable subsystem balancing with limited control authority. We present a learning-based control approach for underactuated balance robots. The tracking and balancing control is designed the controller in fast- and slow-time scales. In the slow-time scale, model predictive control is adopted to plan desired internal state profile to achieve external trajectory tracking task. The internal state is then stabilized around the planned profile in the fast-time scale. The control design is based on a learned Gaussian process (GP) regression model without need of a priori knowledge about the robot dynamics. The controller also incorporates the GP model predicted variance to enhance robustness to modeling errors. Experiments are presented using a Furuta pendulum system. Kuo Chen, Jingang Yi, Dezhen Song |
ICRA | 3 |
| 2019 | Steering Co-centered and Co-directional Optical and Acoustic Beams with a Water-immersible MEMS Scanning Mirror for Underwater Ranging and CommunicationabstractThis paper reports the development of a compact optical-acoustic frontend module for underwater communication and ranging. The module is enabled by a new water-immersible MEMS scanning mirror (WIMSM). It is capable of transmitting, receiving and steering co-centered and co-directional laser and ultrasound beams under water. To monitor its rotating angle in real time, scan position sensors based on Hall effect have been integrated into the WIMSM. The angular alignment of the laser and ultrasound beams in both transmission and reception modes has been examined. The experimental results show that the laser and ultrasound beams can remain aligned with less than 2.1 degrees under envelope of pan and tilt rotations. This capability is critical for the continuing development of the new bi-modal communication and ranging underwater Vehicles (AUVs). Xiaoyu Duan, Dezhen Song |
ICRA | 2 |
| 2019 | Toward Fingertip Non-Contact Material Recognition and Near-Distance Ranging for Robotic GraspingabstractWe report the feasibility study of a new acoustic and optical bi-modal distance & material sensor for robotic grasping. The new sensor is designed to be mounted on the robot fingertip to provide last-moment perception before contact happens. It is based on both pulse-echo ultrasound and optoacoustic effects enabled by single-element air-coupled transducers. In contrast to conventional contact-based and recent pre-touch approaches, this new method overcomes their disadvantages and provides robotic fingers with the capability to detect the distance and material type of the target at a near distance before contact occurs, which is crucial for robust and nimble grasping. The proposed sensor has been tested with different materials, shapes, and porous properties. The experimental results show that this sensor design is functional and practical. Di Wang 0020, Dezhen Song |
ICRA | 3 |
| 2019 | Semi-semantic Line-Cluster Assisted Monocular SLAM for Indoor Environments
Ting Sun 0001, Dezhen Song, Dit-Yan Yeung, Ming Liu 0001 |
ICVS | 2 |
| 2019 | On the Tunable Sparse Graph Solver for Pose Graph Optimization in Visual SLAM ProblemsabstractWe report a tunable sparse optimization solver that can trade a slight decrease in accuracy for significant speed improvement in pose graph optimization in visual simultaneous localization and mapping (vSLAM). The solver is designed for devices with significant computation and power constraints such as mobile phones or tablets. Two approaches have been combined in our design. The first is a graph pruning strategy by exploiting objective function structure to reduce the optimization problem size which further sparsifies the optimization problem. The second step is to accelerate each optimization iteration in solving increments for the gradient-based search in Gauss-Newton type optimization solver. We apply a modified Cholesky factorization and reuse the decomposition result from last iteration by using Cholesky update/downdate to accelerate the computation. We have implemented our solver and tested it with open source data. The experimental results show that our solver can be twice as fast as the counterpart while maintaining a loss of less than 5% in accuracy. Chieh Chou, Di Wang 0020, Dezhen Song, Timothy A. Davis 0001 |
IROS | 3 |
| 2019 | Virtual Lane Boundary Generation for Human-Compatible Autonomous Driving: A Tight Coupling between Perception and PlanningabstractExisting autonomous vehicle (AV) navigation algorithms treat lane recognition, obstacle avoidance, local path planning, and lane following as separate functional modules which result in driving behavior that is incompatible with human drivers. It is imperative to design human-compatible navigation algorithms to ensure transportation safety. We develop a new tightly-coupled perception-planning framework that combines all these functionalities to ensure human-compatibility. Using GPS-camera-lidar sensor fusion, we detect actual lane boundaries (ALBs) and propose availability-reasonability-feasibility (ARF) threefold tests to determine if we should generate virtual lane boundaries (VLBs) or follow ALBs. If needed, VLBs are generated using a dynamically adjustable multi-objective optimization framework that considers obstacle avoidance, trajectory smoothness (to satisfy vehicle kinodynamic constraints), trajectory continuity (to avoid sudden movements), GPS following quality (to execute global plan), and lane following or partial direction following (to meeting human expectation). Consequently, vehicle motion is more human compatible than existing approaches. We have implemented our algorithm and tested under open source data with satisfying results. Binbin Li 0006, Dezhen Song, Ankit Ramchandani, Hsin-Min Cheng, Di Wang 0020, Yiliang Xu, Baifan Chen |
IROS | 2 |
| 2019 | Automatic Pavement Crack Detection by Multi-Scale Image FusionabstractPavement crack detection from images is a challenging problem due to intensity inhomogeneity, topology complexity, low contrast, and noisy texture background. Traditional learning-based approaches have difficulties in obtaining representative training samples. We propose a new unsupervised multi-scale fusion crack detection (MFCD) algorithm that does not require training data. First, we develop a windowed minimal intensity path-based method to extract the candidate cracks in the image at each scale. Second, we find the crack correspondences across different scales. Finally, we develop a crack evaluation model based on a multivariate statistical hypothesis test. Our approach successfully combines strengths from both the large-scale detection (robust but poor in localization) and the small-scale detection (detail-preserving but sensitive to clutter). We analyze and experimentally test the computational complexity of our MFCD algorithm. We have implemented the algorithm and have it extensively tested on three public data sets, including two public pavement data sets and an airport runway data set. Compared with six existing methods, experimental results show that our method outperforms all counterparts. Specifically, it increases the precision, recall, and F1-measure over the state-of-the-art by 22%, 12%, and 19%, respectively, on one public data set. Haifeng Li 0008, Dezhen Song, Binbin Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Encoder-Camera-Ground Penetrating Radar Tri-Sensor Mapping for Surface and Subsurface Transportation Infrastructure InspectionabstractWe report system and algorithmic development for a sensing suite comprising multiple sensors for both surface and subsurface transportation infrastructure inspection focusing on multi-modal mapping for inspection. The sensing suite contains a camera, a ground penetrating radar (GPR), and a wheel encoder. We design the sensing suite and propose a data collection scheme using customized artificial landmarks (ALs). We use ALs to synchronize two types data streams: camera images that are temporally evenly-spaced and GPR/encoder data that are spatially evenly-spaced. We also employ pose graph optimization with synchronization as penalty functions to further refine synchronization and perform data fusion for 3D reconstruction. We have implemented the system and tested it in physical experiments. The results show that our system successfully fuses three sensory data and product metric 3D reconstruction. The sensor fusion approach reduces the end-to-end distance error from 7.45cm to 3.10cm. Chieh Chou, Aaron Kingery, Di Wang 0020, Haifeng Li 0008, Dezhen Song |
ICRA | 5 |
| 2018 | Robotic Subsurface Pipeline Mapping with a Ground-penetrating Radar and a CameraabstractWe propose a novel subsurface pipeline mapping method by fusing Ground Penetrating Radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with non-perpendicular angles. Furthermore, we fuse visual simultaneous localization and mapping outputs, encoder readings with GPR scans to classify hyperbolas into different pipeline groups. We extensively apply the J-Linkage method and maximum likelihood estimation to improve algorithm robustness and accuracy. As the result, we optimally estimate the radii and locations of all pipelines. We have implemented our method and tested it in physical experiments with representative pipeline configurations. The results show that our method successfully reconstructs all subsurface pipes. Moreover, the average localization error is 4.69cm. Haifeng Li 0008, Chieh Chou, Longfei Fan, Binbin Li 0006, Di Wang 0020, Dezhen Song |
IROS | 6 |
| 2018 | Lane Marking Quality Assessment for Autonomous DrivingabstractMeasuring the quality of roads and ensuring they are ready for autonomous driving is important for future transportation systems. Here we focus on developing metrics and algorithms to assess lane marking (LM)qualities from an egocentric view of an inspection vehicle equipped with a global positioning system (GPS)receiver, a frontal-view camera, and a light detection and ranging (LIDAR)system. We propose three quality metrics for LMs: correctness, shape, and visibility. The correctness metric measures the divergence between the expected LMs based on prior map inputs and the actual sensor inputs. The shape metric evaluates smoothness in road curvature and width range. The visibility metric evaluates the contrast between LMs and background road surfaces. We propose a dual-modal algorithm to compute these metrics. We have implemented the algorithms and tested them under KITTI dataset. The results show that our metrics can successfully detect LM anomalies in all testing scenarios. Binbin Li 0006, Dezhen Song, Haifeng Li 0008, Adam Pike, Paul Carlson |
IROS | 2 |
| 2017 | Mirror-assisted calibration of a multi-modal sensing array with a ground penetrating radar and a cameraabstractTo develop a multi-modal in-traffic bridge deck scanning device, we need to estimate the relative pose between a ground penetrating radar (GPR) and a camera. Unlike camera images, GPR output is in a non-Euclidean coordinate system because it only detects underground objects relative to road surface. When road surface is non-planar, its output cannot be trivially mapped to a 3D Cartesian system which is necessary for sensor fusion. Since there is no joint coverage between two sensors due to mounting requirements, we design an artificial planar bridge assisted by a planar mirror as the calibration rig. We combine the pinhole camera model with mirror reflection transformation and model the GPR imaging process. We estimate the camera and mirror poses and extract readings from hyperbolas generated from metal balls. We employ the maximum likelihood estimator to estimate the rigid body transformation between the two sensors and provide the closed form error analysis. We have conducted physical experiments to validate our calibration process and shown the average error of 6.67 mm for our calibration model. The result is satisfying considering the GPR signal wave length is 18.75 cm. Chieh Chou, Shu-Hao Yeh, Dezhen Song |
IROS | 3 |
| 2017 | Localization in Inconsistent WiFi Environments
Hsin-Min Cheng, Dezhen Song |
ISRR | 2 |
| 2017 | Probabilistic Boundary Coverage for Unknown Target Fields with Large Perception Uncertainty and Limited Sensing Range
Binbin Li 0006, Dezhen Song |
ISRR | 2 |
| 2017 | Sharing Heterogeneous Spatial Knowledge: Map Fusion Between Asynchronous Monocular Vision and Lidar or Other Prior Inputs
Joseph Lee, Shu-Hao Yeh, Hsin-Min Cheng, Baifan Chen, Dezhen Song |
ISRR | 6 |
| 2016 | Visual programming for mobile robot navigation using high-level landmarksabstractWe propose a visual programming system that allows users to specify navigation tasks for mobile robots using high-level landmarks in a virtual reality (VR) environment constructed from the output of visual simultaneous localization and mapping (vSLAM). The VR environment provides a Google Street View-like interface for users to familiarize themselves with the robot's working environment, specify high-level landmarks, and determine task-level motion commands related to each landmark. Our system builds a roadmap by using the pose graph from the vSLAM outputs. Based on the roadmap, the high-level landmarks, and task-level motion commands, our system generates an output path for the robot to accomplish the navigation task. We present data structures, architecture, interface, and algorithms for our system and show that, given nssearch-type motion commands, our system generates a path in O(ns(nrlognr+mr)) time, where nrand mrare the number of roadmap nodes and edges, respectively. We have implemented our system and tested it on real world data. Joseph Lee, Yiliang Xu, Dezhen Song |
IROS | 4 |
| 2015 | Robust RGB-D Odometry Using Point and Line FeaturesabstractLighting variation and uneven feature distribution are main challenges for indoor RGB-D visual odometry where color information is often combined with depth information. To meet the challenges, we fuse point and line features to form a robust odometry algorithm. Line features are abundant indoors and less sensitive to lighting change than points. We extract 3D points and lines from RGB-D data, analyze their measurement uncertainties, and compute camera motion using maximum likelihood estimation. We prove that fusing points and lines produces smaller motion estimate uncertainty than using either feature type alone. In experiments we compare our method with state-of-the-art methods including a keypoint-based approach and a dense visual odometry algorithm. Our method outperforms the counterparts under both constant and varying lighting conditions. Specifically, our method achieves an average translational error that is 34.9% smaller than the counterparts, when tested using public datasets. Dezhen Song |
ICCV | 2 |
| 2015 | A robotic bipedal model for human walking with slipsabstractSlip is the major cause of falls in human locomotion. We present a new bipedal modeling approach to capture and predict human walking locomotion with slips. Compared with the existing bipedal models, the proposed slip walking model includes the human foot rolling effects, the existence of the double-stance gait and active ankle joints. One of the major developments is the relaxation of the nonslip assumption that is used in the existing bipedal models. We conduct extensive experiments to optimize the gait profile parameters and to validate the proposed walking model with slips. The experimental results demonstrate that the model successfully predicts the human recovery gaits with slips. Kuo Chen, Mitja Trkov, Jingang Yi, Yizhai Zhang, Tao Liu 0006, Dezhen Song |
ICRA | 6 |
| 2015 | Robustness to lighting variations: An RGB-D indoor visual odometry using line segmentsabstractLarge lighting variation challenges all visual odometry methods, even with RGB-D cameras. Here we propose a line segment-based RGB-D indoor odometry algorithm robust to lighting variation. We know line segments are abundant indoors and less sensitive to lighting change than point features. However, depth data are often noisy, corrupted or even missing for line segments which are often found on object boundaries where significant depth discontinuities occur. Our algorithm samples depth data along line segments, and uses a random sample consensus approach to identify correct depth and estimate 3D line segments. We analyze 3D line segment uncertainties and estimate camera motion by minimizing the Mahalanobis distance. In experiments we compare our method with two state-of-the-art methods including a keypoint-based approach and a dense visual odometry algorithm, under both constant and varying lighting. Our method demonstrates superior robustness to lighting change by outperforming the competing methods on 6 out of 8 long indoor sequences under varying lighting. Meanwhile our method also achieves improved accuracy even under constant lighting when tested using public data. Dezhen Song |
IROS | 2 |
| 2015 | Automatic Bird Species Filtering Using a Multimodel ApproachabstractWe report a filtering algorithm for bird species detection using videos captured by uncalibrated moving cameras, a typical characteristic of crowd sourced videos. The algorithm tracks both body and wing motion dimensions of a flying bird to form signatures for species filtering. In the body model, we consider both cases when background motion introduced by the camera can and cannot be directly recognized using key point matching. We are also able to recover intrinsic camera parameters in the body motion tracking. In the wing model, we consider both periodic wing flapping and gliding motion patterns. These models are combined to form a multimodel framework. We have tested the algorithm and compared its performance with single model approaches in physical experiments. Results show that the new algorithm significantly reduces the false positive rate, while maintaining a low false negative rate. The area under ROC curve is 92.86%. Dezhen Song |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2015 | Visual Navigation Using Heterogeneous Landmarks and Unsupervised Geometric ConstraintsabstractWe present a heterogeneous landmark-based visual navigation approach for a monocular mobile robot. We utilize heterogeneous visual features, such as points, line segments, lines, planes, and vanishing points, and their inner geometric constraints managed by a novel multilayer feature graph (MFG). Our method extends the local bundle adjustment-based visual simultaneous localization and mapping (SLAM) framework by explicitly exploiting the heterogeneous features and their inner geometric relationships in an unsupervised manner. As the result, our heterogeneous landmark-based visual navigation algorithm takes a video stream as input, initializes and iteratively updates MFG based on extracted key frames, and refines robot localization and MFG landmarks through the process. We present pseudocode for the algorithm and analyze its complexity. We have evaluated our method and compared it with state-of-the-art point landmark-based visual SLAM methods using multiple indoor and outdoor datasets. In particular, on the KITTI dataset, our method reduces the translational error by 52.5% under urban sequences where rectilinear structures dominate the scene. Dezhen Song |
IEEE Trans. Robotics | 2 |
| 2014 | An Efficient Online Hierarchical Supervoxel Segmentation Algorithm for Time-critical Applications
Yiliang Xu, Dezhen Song, Anthony Hoogs |
BMVC | 2 |
| 2014 | Toward featureless visual navigation: Simultaneous localization and planar surface extraction using motion vectors in video streamsabstractUnlike the traditional feature-based methods, we propose using motion vectors (MVs) from video streams as inputs for visual navigation. Although MVs are very noisy and with low spatial resolution, MVs do possess high temporal resolution which means it is possible to merge MVs from different frames to improve signal quality. Homography filtering and MV thresholding are proposed to further improve MV quality so that we can establish plane observations from MVs. We propose an extended Kalman filter (EKF) based approach to simultaneously track robot motion and planes. We formally model error propagation of MVs and derive variance of the merged MVs. We have implemented the proposed method and tested it in physical experiments. Results show that the system is capable of performing robot localization and plane mapping with a relative trajectory error of less than 5.1%. Dezhen Song |
ICRA | 2 |
| 2014 | High level landmark-based visual navigation using unsupervised geometric constraints in local bundle adjustmentabstractWe present a high level landmark-based visual navigation approach for a monocular mobile robot. We utilize heterogeneous features, such as points, line segments, lines, planes, and vanishing points, and their inner geometric constraints as the integrated high level landmarks. This is managed through a multilayer feature graph (MFG). Our method extends local bundle adjustment (LBA)-based framework by explicitly exploiting different features and their geometric relationships in an unsupervised manner. The algorithm takes a video stream as input, initializes and incrementally updates MFG based on extracted key frames; it also refines localization and MFG landmarks through the LBA. Physical experiments show that our method can reduce the absolute trajectory error of a traditional point landmark-based LBA method by up to 63.9%. Dezhen Song, Jingang Yi |
ICRA | 2 |
| 2014 | Stationary balance control of a bikebotabstractWe present the development of the gyroscopic-balanced control of an autonomous bikebot. The bikebot is an actively controlled bicycle-based robotic platform with a gyro-balancer developed to study human dynamic postural balance motor skills through unstable physical human-robot interactions. We also present a dynamic model and analysis for stationary bikebot. A nonlinear balancing controller is designed to stabilize the underactuated stationary bikebot on an orbital trajectory around the unstable equilibrium point that is coupled with another orbit of the actuated gyro-balancer. We then demonstrate the analysis and control design with experimental validations. Finally, we present a set of human riding experiments to show how the bikebot can be used to perturb and excite human sensorimotor feedback loop for dynamic postural balance motor skills. Yizhai Zhang, Pengcheng Wang 0002, Jingang Yi, Dezhen Song, Tao Liu 0006 |
ICRA | 4 |
| 2014 | Planar building facade segmentation and mapping using appearance and geometric constraintsabstractSegmentation and mapping of planar building facades (PBFs) can increase a robot's ability of scene understanding and localization in urban environments which are often quasi-rectilinear and GPS-challenged. PBFs are basic components of the quasi-rectilinear environment. We propose a passive vision-based PBF segmentation and mapping algorithm by combining both appearance and geometric constraints. We propose a rectilinear index which allows us to segment out planar regions using appearance data. Then we combine geometric constraints such as reprojection errors, orientation constraints, and coplanarity constraints in an optimization process to improve the mapping of PBFs. We have implemented the algorithm and tested it in comparison with state-of-the-art. The results show that our method can reduce the angular error of scene structure by an average of 82.82%. Joseph Lee, Dezhen Song |
IROS | 3 |
| 2014 | Featureless Motion Vector-Based Simultaneous Localization, Planar Surface Extraction, and Moving Obstacle Tracking
Dezhen Song |
WAFR | 2 |
| 2014 | Automatic Bird Species Detection From Crowd Sourced VideosabstractTo assist nature observation, we develop two algorithms to enable automatic bird species filtering using crowd sourced videos as inputs where camera motion and parameters are often unknown. The first algorithm recognizes the time series of salient extremities, which is the inter-wing tip distance (IWTD), from motion segmented bird contours. To analyze the feasibility of the proposed algorithm, we derive the probability that the salient extremity can be recognized from a video captured by an arbitrary camera with unknown parameters. We also prove that the periodicity of the IWTD in the image is the same as the wingbeat frequency in the 3D space regardless of camera parameters with the exception of ignorable degenerated cases. Therefore, the second algorithm applies Fast Fourier Transform to the series and classifies bird species using likelihood ratios. The algorithm outputs a ranked list of likelihood of candidate species. Experiment results validate our analysis and show that the algorithm is very robust to segmentation error and data loss up to 30%. Dezhen Song |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | Cooperative Search of Multiple Unknown Transient Radio Sources Using Multiple Paired Mobile RobotsabstractWe develop a localization method to enable a team of mobile robots to search for multiple unknown transient radio sources. Because of signal source anonymity, short transmission durations, and dynamic transmission patterns, robots cannot treat the radio sources as continuous radio beacons. Moreover, robots do not know the source transmission power and have limited sensing ranges. To cope with these challenges, we pair up robots and develop a cooperative sensing model using signal strength ratios from the paired robots. We formally prove that the joint conditional posterior probability of source locations for the m-robot team can be obtained by combining the pairwise joint posterior probabilities, which can be derived from signal strength ratios. Moreover, we propose a pairwise ridge walking algorithm (PRWA) to coordinate the robot pairs based on the clustering of high-probability regions and the minimization of local Shannon entropy. We have implemented and validated the algorithm under both the hardware-driven simulation and physical experiments. Experimental results show that the PRWA-based localization scheme consistently outperforms the other four heuristics. Chang-Young Kim, Dezhen Song, Yiliang Xu, Jingang Yi, Xinyu Wu 0001 |
IEEE Trans. Robotics | 2 |
| 2013 | Hierarchical activity discovery within spatio-temporal context for video anomaly detectionabstractIn this paper, we present a novel approach for video anomaly detection in crowded and complicated scenes. The proposed approach detects anomalies based on a hierarchical activity pattern discovery framework comprehensively considering both global and local spatio-temporal contexts. The discovery is a coarse-to-fine learning process with unsupervised ways for automatically constructing normal activity patterns at different levels. An unified anomaly energy function is designed based on these discovered activity patterns to identify the abnormal level of an input motion pattern. We demonstrate the efficiency of the proposed method on the UCSD anomaly detection datasets (Ped1 and Ped2) and compare the performance with existing work. Dan Xu 0006, Xinyu Wu 0001, Dezhen Song, Nannan Li 0001, Yen-Lun Chen |
ICIP | 3 |
| 2013 | Decentralized searching of multiple unknown and transient radio sourcesabstractWe develop a decentralized algorithm to coordinate a group of mobile robots to search for unknown and transient radio sources. In addition to limited mobility and ranges of communication and sensing, the robot team has to deal with challenges from signal source anonymity, short transmission duration, and variable transmission power. We propose a two-step approach: first, we decentralize belief functions that robots use to track source locations using checkpoint-based synchronization, and second, we propose a decentralized planning strategy to coordinate robots to ensure the existence of checkpoints. We analyze memory usage, data amount in communication, and searching time for the proposed algorithm. We have implemented the proposed algorithm and compared it with two heuristics. The experiment results show that our algorithm successfully trades a modest amount of memory for the fastest searching time among the three methods. Chang-Young Kim, Dezhen Song, Jingang Yi |
ICRA | 2 |
| 2013 | Automatic bird species detection using periodicity of salient extremitiesabstractTo assist nature observation, we develop an automatic bird species filtering method that takes videos from cameras with unknown parameters as input, and outputs likelihood of candidate species. The method recognizes the time series of salient extremities, which is the inter-wing tip distance, performs frequency analysis on periodicity, and provides a species prediction metric using likelihood ratios. To analyze the feasibility of the proposed method, we derive the probability that the salient extremity can be recognized in image for an arbitrary camera perspective.We also prove that the periodicity of the IWTD in the image is the same as the wingbeat frequency in the 3D space regardless of camera parameters with the exception of ignorable degenerated cases. Experiment results validate our analysis and show that the algorithm is very robust to segmentation error and data loss up to 30%. Dezhen Song |
ICRA | 2 |
| 2012 | A two-view based multilayer feature graph for robot navigationabstractTo facilitate scene understanding and robot navigation in a modern urban area, we design a multilayer feature graph (MFG) based on two views from an on-board camera. The nodes of an MFG are features such as scale invariant feature transformation (SIFT) feature points, line segments, lines, and planes while edges of the MFG represent different geometric relationships such as adjacency, parallelism, collinearity, and coplanarity. MFG also connects the features in two views and the corresponding 3D coordinate system. Building on SIFT feature points and line segments, MFG is constructed using feature fusion which incrementally, iteratively, and extensively verifies the aforementioned geometric relationships using random sample consensus (RANSAC) framework. Physical experiments show that MFG can be successfully constructed in urban area and the construction method is demonstrated to be very robust in identifying feature correspondence. Haifeng Li 0008, Dezhen Song, Jingtai Liu |
ICRA | 2 |
| 2012 | Path planning for clothes climbing robots on deformable clothes surfaceabstractThis paper proposes a novel path planning method for a robot to climb on the deformable clothes surface. Based on the deformable characteristic of the clothes, the tension force of clothes is analyzed and the model of tension degree is established. A clothes climbing robot called Clothbot is composed of a two-wheeled gripper and a 2 Degrees of Freedom (DOF) tail. Based on the locomotion of this robot, the weights of tension degree and the locomotion characteristic are added into the A* algorithm. Combined with the two weights applied, the optimal path to the target for the Clothbot is obtained. The Clothbot has been developed to evaluate the algorithm. The simulation and the experiments have verified the feasibility of this method. In addition, The error state of the movement of the robot which is called side tumbling has been corrected by the motion of the 2-DOF tail. Xinyu Wu 0001, Dezhen Song, Ruiqing Fu, Duan Zheng, Yangsheng Xu |
IROS | 3 |
| 2012 | Simultaneous Localization of Multiple Unknown and Transient Radio Sources Using a Mobile RobotabstractWe report system and algorithm developments that utilize a single mobile robot to simultaneously localize multiple unknown transient radio sources. Because of signal source anonymity, short transmission durations, and dynamic transmission patterns, the robot cannot treat the radio sources as continuous radio beacons. To deal with this challenging localization problem, we model the radio source behaviors using a novel spatiotemporal probability occupancy grid that captures transient characteristics of radio transmissions and tracks posterior probability distributions of radio sources. As a Monte Carlo method, a ridge walking motion planning algorithm is proposed to enable the robot to efficiently traverse the high-probability regions to accelerate the convergence of the posterior probability distribution. We also formally show that the time to find a radio source is insensitive to the number of radio sources, and hence, our algorithm has great scalability. We have implemented the algorithms and extensively tested them in comparison with two heuristic methods: a random walk and a fixed-route patrol. The localization time of our algorithms is consistently shorter than that of the two heuristic methods. Dezhen Song, Chang-Young Kim, Jingang Yi |
IEEE Trans. Robotics | 1 |
| 2011 | Robust recognition of planar mirrored walls using a single viewabstractWe report a method for the detection and recognition of a large planar mirror based on the images captured by a monocular camera. We start with deriving a mirror transformation matrix in a homogeneous coordinate and geometric constraints for corresponding real and virtual feature points in the image. We find that existing feature detection methods are not reflection invariant. We introduce a secondary artificial reflection to virtual features to generate secondary features which are proven to share a rigid body motion relationship with the original feature set. We propose an iterative strategy to adjust the secondary mirror configuration so that existing feature matching methods can be used. The combined method yields a robust mirror detection algorithm which has been verified in physical experiments. Ali-akbar Agha-mohammadi, Dezhen Song |
ICRA | 2 |
| 2011 | Localization of multiple unknown transient radio sources using multiple paired mobile robots with limited sensing rangesabstractWe develop a localization method enabling a team of mobile robots to search for multiple unknown transient radio sources. Due to signal source anonymity, short transmission durations, and dynamic transmission patterns, robots cannot treat the radio sources as continuous radio beacons. Moreover, robots do not know the source transmission power and have limited sensing ranges. To cope with these challenges, we pair up robots and develop a sensing model using the signal strength ratio from the paired robots. We formally prove that the sensed conditional joint posterior probability of source locations for the m-robot team can be obtained by combining the pairwise joint posterior probabilities, which can be derived from signal strength ratios. Moreover, we propose a pairwise ridge walking algorithm (PRWA) to coordinate the robot pairs based on the clustering of high probability regions and the minimization of local Shannon entropy. We have implemented and validated the algorithm under hardware-driven simulation. Chang-Young Kim, Dezhen Song, Yiliang Xu, Jingang Yi |
ICRA | 2 |
| 2011 | Balance control and analysis of stationary riderless motorcyclesabstractWe present balancing control analysis of a stationary riderless motorcycle. We first present the motorcycle dynamics with an accurate steering mechanism model with consideration of lateral movement of the tire/ground contact point. A nonlinear balance controller is then designed. We estimate the domain of attraction (DOA) of motorcycle dynamics under which the stationary motorcycle can be stabilized by steering. For a typical motorcycle/bicycle configuration, we find that the DOA is relatively small and thus balancing control by only steering at stationary is challenging. The balance control and DOA estimation schemes are validated by experiments conducted on the Rutgers autonomous motorcycle. The attitudes of the motorcycle platform are obtained by a novel estimation scheme that fuses measurements from global positioning systems (GPS) and inertial measurement units (IMU). We also present the experiments of the GPS/IMU-based attitude estimation scheme in the paper. Yizhai Zhang, Jingliang Li, Jingang Yi, Dezhen Song |
ICRA | 4 |
| 2011 | Optimal Scheduling of Multicluster Tools With Constant Robot Moving Times, Part II: Tree-Like Topology ConfigurationsabstractIn this paper, we analyze optimal scheduling of a tree-like multicluster tool with single-blade robots and constant robot moving times. We present a recursive minimal cycle time algorithm to reveal a multi-unit resource cycle for multicluster tools under a given robot schedule. For a serial-cluster tool, we provide a closed-form formulation for the minimal cycle time. The formulation explicitly provides the interaction relationship among clusters. We further present decomposition conditions under which the optimal scheduling of multicluster becomes much easier and straightforward. Optimality conditions for the widely used robot pull schedule are also provided. An example from industry production is used to illustrate the analytical results. The decomposition and optimality conditions for the robot pull schedule are also illustrated by Monte Carlo simulation for the industrial example. Wai Kin Chan, Shengwei Ding, Jingang Yi, Dezhen Song |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2011 | On the Time to Search for an Intermittent Signal Source Under a Limited Sensing RangeabstractA mobile robot with a limited sensing range is deployed to search for a stationary target that intermittently emits short duration signals. The searching mission is accomplished as soon as the robot receives a signal from the target. We propose the expected searching time (EST) as a primary metric to evaluate different robot motion plans under different robot configurations. To illustrate the proposed metric, we present two case studies. In the first case, we analyze two common motion plans: a slap method (SM) and a random walk (RW). The EST analysis shows that the SM is asymptotically faster than the RW when the searching space size increases. In the second case, we compare a team ofnhomogeneous low-cost robots with a super robot that has the sensing range equal to that of the summation of thenrobots. Our analysis shows that the low-cost robot team takes Θ(1/n) time, while the super robot takes Θ(1/√n) time asn→ ∞. Our metrics successfully demonstrate their ability in assessing the searching performance. The analytical results are also confirmed in simulation and physical experiments. Dezhen Song, Chang-Young Kim, Jingang Yi |
IEEE Trans. Robotics | 1 |
| 2010 | A Low False Negative Filter for Detecting Rare Bird Species from Short Video Segments using a Probable Observation Data Set-based EKF MethodabstractWe report a new filter for assisting the search for rare bird species. Since a rare bird only appears in front of the camera with very low occurrence (e.g. less than ten times per year) for very short duration (e.g. less than a fraction of a second), our algorithm must have very low false negative rate. We verify the bird body axis information with the known bird flying dynamics from the short video segment. Since a regular extended Kalman filter (EKF) cannot converge due to high measurement error and limited data, we develop a novel Probable Observation Data Set (PODS)-based EKF method. The new PODS-EKF searches the measurement error range for all probable observation data that ensures the convergence of the corresponding EKF in short time frame. The algorithm has been extensively tested in experiments. The results show that the algorithm achieves 95.0% area under ROC curve in physical experiment with close to zero false negative rate. Dezhen Song, Yiliang Xu |
AAAI | 1 |
| 2010 | Error Aware Monocular Visual Odometry using Vertical Line Pairs for Small Robots in Urban AreasabstractWe report a new error-aware monocular visual odometry method that only uses vertical lines, such as vertical edges of buildings and poles in urban areas as landmarks. Since vertical lines are easy to extract, insensitive to lighting conditions/ shadows, and sensitive to robot movements on the ground plane, they are robust features if compared with regular point features or line features. We derive a recursive visual odometry method based on the vertical line pairs. We analyze how errors are propagated and introduced in the continuous odometry process by deriving the closed form representation of covariance matrix. We formulate the minimum variance ego-motion estimation problem and present a method that outputs weights for different vertical line pairs. The resulting visual odometry method is tested in physical experiments and compared with two existing methods that are based on point features and line features, respectively. The experiment results show that our method outperforms its two counterparts in robustness, accuracy, and speed. The relative errors of our method are less than 2% in experiments. Dezhen Song |
AAAI | 2 |
| 2010 | A Low False Negative Filter for Detecting Rare Bird Species From Short Video Segments Using a Probable Observation Data Set-Based EKF MethodabstractWe report a new filter to assist the search for rare bird species. Since a rare bird only appears in front of a camera with very low occurrence (e.g., less than ten times per year) for very short duration (e.g., less than a fraction of a second), our algorithm must have a very low false negative rate. We verify the bird body axis information with the known bird flying dynamics from the short video segment. Since a regular extended Kalman filter (EKF) cannot converge due to high measurement error and limited data, we develop a novel probable observation data set (PODS)-based EKF method. The new PODS-EKF searches the measurement error range for all probable observation data that ensures the convergence of the corresponding EKF in short time frame. The algorithm has been extensively tested using both simulated inputs and real video data of four representative bird species. In the physical experiments, our algorithm has been tested on rock pigeons and red-tailed hawks with 119 motion sequences. The area under the ROC curve is 95.0%. During the one-year search of ivory-billed woodpeckers, the system reduces the raw video data of 29.41 TB to only 146.7 MB (reduction rate 99.9995%). Dezhen Song, Yiliang Xu |
IEEE Trans. Image Process. | 1 |
| 2009 | Monte Carlo simultaneous localization of multiple unknown transient radio sources using a mobile robot with a directional antennaabstractWe report our system and algorithm developments that enable a single mobile robot equipped with a directional antenna to simultaneously localize multiple unknown transient radio sources. Due to signal source anonymity, short transmission durations, and dynamic transmission patterns, the robot cannot treat the radio sources as continuous radio beacons.We model the radio source behaviors using a novel spatiotemporal probability occupancy grid (SPOG) that captures transient characteristics of radio transmissions and tracks the spatiotemporal posterior probability distribution of the radio transmissions. As a Monte Carlo method, we propose a ridge walking motion planning algorithm that enables the robot to efficiently traverse the high probability regions to accelerate the convergence of the posterior probability distribution. We have implemented the algorithms and the experiment results show that our method consistently outperforms methods such as a random walk or a fixed-route patrol mechanism. Dezhen Song, Chang-Young Kim, Jingang Yi |
ICRA | 1 |
| 2009 | Modeling and motion stability analysis of skid-steered mobile robotsabstractSkid-steered mobile robots are widely used because of the simplicity of mechanism and high reliability. However, understanding of the kinematics and dynamics of such a robotic platform is challenging due to the complex wheel/ground interactions and kinematic constraints. In this paper, we attempt to develop a kinematic and dynamic modeling scheme to analyze the skid-steered mobile robot. We model wheel/ground interaction and analyze the robot motion stability. As an application example, we present how to utilize the kinematic and dynamic modeling and analysis for robot localization and slip estimation using only low-cost strapdown inertial measurement units (IMU). The extended Kalman filter (EKF)-based localization scheme incorporates the kinematic constraints. The performance of the EKF-based localization and slip estimation scheme are presented. The estimation methodology is tested and validated on a robotic testbed. Jingang Yi, Dezhen Song, Suhada Jayasuriya, Jingtai Liu |
ICRA | 4 |
| 2009 | Systems and algorithms for autonomously simultaneous observation of multiple objects using robotic PTZ cameras assisted by a wide-angle cameraabstractWe report an autonomous observation system with multiple pan-tilt-zoom (PTZ) cameras assisted by a fixed wide-angle camera. The wide-angle camera provides large but low resolution coverage and detects and tracks all moving objects in the scene. Based on the output of the wide-angle camera, the system generates spatiotemporal observation requests for each moving object, which are candidates for close-up views using PTZ cameras. Due to the fact that there are usually much more objects than the number of PTZ cameras, the system first assigns a subset of the requests/objects to each PTZ camera. The PTZ cameras then select the parameter settings that best satisfy the assigned competing requests to provide high resolution views of the moving objects. We solve the request assignment and the camera parameter selection problems in real time. The effectiveness of the proposed system is validated in comparison with an existing work using simulation. The simulation results show that in heavy traffic scenarios, our algorithm increases the number of observed objects by over 200%. Yiliang Xu, Dezhen Song |
IROS | 2 |
| 2009 | On the error analysis of vertical line pair-based monocular visual odometry in urban areaabstractWhen a robot travels in urban area, Global Positional System (GPS) signals might be obstructed by buildings. Hence visual odometry is a choice. We notice that the vertical edges from high buildings and poles of street lights are a very stable set of features that can be easily extracted. Thus, we develop a monocular vision-based odometry system that utilizes the vertical edges from the scene to estimate the robot ego-motion. Since it only takes a single vertical line pair to estimate the robot ego-motion on the road plane, here we model the ego-motion estimation process and analyze how the choice of different vertical line pair impacts the accuracy of the ego-motion estimation process. The resulting closed form error model can assist to choose an appropriate pair of vertical lines to reduce the error in computation. We have implemented the proposed method and validated the error analysis results in physical experiments. Dezhen Song |
IROS | 2 |
| 2009 | Kinematic Modeling and Analysis of Skid-Steered Mobile Robots With Applications to Low-Cost Inertial-Measurement-Unit-Based Motion EstimationabstractSkid-steered mobile robots are widely used because of their simple mechanism and high reliability. Understanding the kinematics and dynamics of such a robotic platform is, however, challenging due to the complex wheel/ground interactions and kinematic constraints. In this paper, we develop a kinematic modeling scheme to analyze the skid-steered mobile robot. Based on the analysis of the kinematics of the skid-steered mobile robot, we reveal the underlying geometric and kinematic relationships between the wheel slips and locations of the instantaneous rotation centers. As an application example, we also present how to utilize the modeling and analysis for robot positioning and wheel slip estimation using only low-cost strapdown inertial measurement units. The robot positioning and wheel slip-estimation scheme is based on an extended Kalman filter (EKF) design that incorporates the kinematic constraints for accuracy enhancement. The performance of the EKF-based positioning and wheel slip-estimation scheme are also presented. The estimation methodology is tested and validated experimentally on a robotic test bed. Jingang Yi, Dezhen Song, Suhada Jayasuriya, Jingtai Liu |
IEEE Trans. Robotics | 4 |
| 2008 | An approximation algorithm for the least overlapping p-Frame problem with non-partial coverage for networked robotic camerasabstractWe report our algorithmic development of the pframe problem that addresses the need of coordinating a set of p networked robotic pan-tilt-zoom cameras for n, (n ≫ p), competing polygonal requests. We assume that the p frames have almost no overlap on the coverage between frames and a request is satisfied only if it is fully covered. We then propose a Resolution Ratio with Non-Partial Coverage (RRNPC) metric to quantify the satisfaction level for a given request with respect to a set of p candidate frames. We propose a latticebased approximation algorithm to search for the solution that maximizes the overall satisfaction. The algorithm builds on an induction-like approach that finds the relationship between the solution to the (p — 1)-frame problem and the solution to the p-frame problem. For a given approximation bound ε, the algorithm runs in O(n/ε3+p2/ε6) time. We have implemented the algorithm and experimental results are consistent with our complexity analysis. Yiliang Xu, Dezhen Song, Jingang Yi, A. Frank van der Stappen |
ICRA | 2 |
| 2008 | On the Analysis of the Depth Error on the Road Plane for Monocular Vision-Based Robot Navigation
Dezhen Song, Hyun Nam Lee, Jingang Yi |
WAFR | 1 |
| 2008 | Steady-State Throughput and Scheduling Analysis of Multicluster Tools: A Decomposition ApproachabstractCluster tools are widely used as semiconductor manufacturing equipment. While throughput analysis and scheduling of single-cluster tools have been well-studied, research work on multicluster tools is still at an early stage. In this paper, we analyze steady-state throughput and scheduling of multicluster tools. We consider the case where all wafers follow the same visit flow within a multicluster tool. We propose a decomposition method that reduces a multicluster tool problem to multiple independent single-cluster tool problems. We then apply the existing and extended results of throughput and scheduling analysis for each single-cluster tool. Computation of lower-bound cycle time (fundamental period) is presented. Optimality conditions and robot schedules that realize such lower-bound values are then provided using ldquopullrdquo and ldquoswaprdquo strategies for single-blade and double-blade robots, respectively. For an -cluster tool, we present lower-bound cycle time computation and robot scheduling algorithms. The impact of buffer/process modules on throughput and robot schedules is also studied. A chemical vapor deposition tool is used as an example of multicluster tools to illustrate the decomposition method and algorithms. The numerical and experimental results demonstrate that the proposed decomposition approach provides a powerful method to analyze the throughput and robot schedules of multicluster tools. Jingang Yi, Shengwei Ding, Dezhen Song, Mike Tao Zhang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2007 | Scheduling Analysis of Cluster Tools with Buffer/Process ModulesabstractModeling and scheduling of cluster tools are critical to improving the productivity and to enhancing the design of wafer processing flows and equipment for semiconductor manufacturing. In this paper, we extend the decomposition methods in the work of Dwande et al. (2005) for multi-cluster tools with buffer/process modules (BPMs). The computation of the lower-bound cycle time (fundamental period) is presented. Optimality conditions and robot schedules that realize such lower-bound values are then provided using "pull" and "swap" strategies for single-blade and double-blade robots, respectively. The impact of BPMs on throughput and robot schedules is studied. It is found that such an impact depends on the BPM processing time and the cycle times of the decomposed clusters on both sides of BPMs. A chemical vapor deposition (CVD) tool is used as an example of multi-cluster tools to illustrate the proposed method, analysis, and algorithms. The numerical and experimental results demonstrate the effectiveness and efficiency of the algorithms. Jingang Yi, Shengwei Ding, Dezhen Song, Mike Tao Zhang |
ICRA | 3 |
| 2007 | Adaptive Trajectory Tracking Control of Skid-Steered Mobile RobotsabstractSkid-steered mobile robots have been widely used for terrain exploration and navigation. In this paper, we present an adaptive trajectory control design for a skid-steered wheeled mobile robot. Kinematic and dynamic modeling of the robot is first presented. A pseudo-static friction model is used to capture the interaction between the wheels and the ground. An adaptive control algorithm is designed to simultaneously estimate the wheel/ground contact friction information and control the mobile robot to follow a desired trajectory. A Lyapunov-based convergence analysis of the controller and the estimation of the friction model parameter are presented. Simulation and preliminary experimental results based on a four-wheel robot prototype are demonstrated for the effectiveness and efficiency of the proposed modeling and control scheme Jingang Yi, Dezhen Song, Zane Goodwin |
ICRA | 2 |
| 2007 | On-demand sharing of a high-resolution panorama video from networked robotic camerasabstractDue to their flexibility in coverage and resolution, networked robotic cameras become more and more popular in applications such as natural observation, security surveillance, and distance learning. Equipped with a high optical zoom lens, a networked robotic camera can generate a giga-pixel-level panorama to cover its viewable region. As new live frames enter the system, this panorama can be updated as a panorama video. User requests are usually not limited to the current camera frame. A user may request a specific region associated with a specific time window. To satisfy different spatiotemporal requests for multiple concurrent users, we present systems and algorithms to allow the on-demand sharing of the high-resolution panorama video. The high-resolution panorama video is encoded into a patch-based representation to allow efficient storage and on-demand content delivery. We present system architecture, user interface, data representation, and encoding/decoding algorithms. In the experiment, we have implemented the system using the MPEG-2 codec. Experimental results show that our system can not only satisfy different spatiotemporal queries but also significantly reduce computation time and communication bandwidth requirement. Ni Qin, Dezhen Song |
IROS | 2 |
| 2007 | IMU-based localization and slip estimation for skid-steered mobile robotsabstractLocalization and wheel slip estimation of a skid- steered mobile robot is challenging because of the complex wheel/ground interactions and kinematics constraints. In this paper, we present a localization and slip estimation scheme for a skid-steered mobile robot using low-cost inertial measurement units (IMU). We first analyze the kinematics of the skid-steered mobile robot and present a nonlinear Kalman filter (KF)- based simultaneous localization and slip estimation scheme. The KF-based localization design incorporates the wheel slip estimation and utilizes robot velocity constraints and estimates to overcome the large drift resulting from the integration of the IMU acceleration measurements. The estimation methodology is tested and validated experimentally with a computer vision- based localization system. Jingang Yi, Dezhen Song, Suhada Jayasuriya |
IROS | 3 |
| 2007 | Approximate Algorithms for a Collaboratively Controlled Robotic CameraabstractDeployed as a natural environment observatory or a surveillance device, a remote networked robotic pan-tilt-zoom camera needs to be controlled by simultaneous frame requests from both online users and in situ sensors such as motion detectors. This paper presents algorithms that are capable of finding a camera frame that optimizes a measure of total satisfaction over all requests, which is a generalized version of the single frame-selection problem proposed by Song et al. in 2006.We present a lattice-based approximation algorithm; given n requests and approximation bound ∈, we analyze the tradeoff between solution quality and the corresponding computation time, and prove that the algorithm runs in O(n/∈3) time. We also develop a branch-and-bound-like implementation that reduces the constant factor of the algorithm by more than 70%. We have implemented the algorithms, and numerical experiment results conform to our analysis. Field experiments of the proposed algorithms have been conducted in the past three years. The proposed algorithms have been deployed successfully in a variety of real world applications including natural environment observation, building construction monitoring, and the surveillance of public space. Dezhen Song, Kenneth Y. Goldberg |
IEEE Trans. Robotics | 1 |
| 2006 | Aligning Windows of Live Video From an Imprecise Pan-tilt-zoom Robotic Camera into a Remote Panoramic DisplayabstractA pan-tilt-zoom robotic camera can provide detailed live video of selected areas of interest within a large potential viewing field. To provide spatial context for human observers, it is desirable to insert the resulting live video into a large spherical panoramic display representing the entire viewing field. Accurate alignment of the video stream within the panoramic display is difficult due to small errors in the robot pan-tilt values and image distortion due to nonlinear projection. Existing image alignment algorithms cannot keep up with rapid changes in camera position. In this paper, we present a constant-time image alignment algorithm based on spherical projection and projection-invariant selective sampling that accurately registers paired images at 25 frames per second on a standard PC. Experiments suggest that the new alignment algorithm is faster than previous algorithms by a factor four or more. In a companion paper, we present a new calibration algorithm based on image variance density that optimally estimates camera pan-tilt parameters Ni Qin, Dezhen Song, Kenneth Y. Goldberg |
ICRA | 2 |
| 2006 | A Minimum Variance Calibration Algorithm for Pan-tilt Robotic Cameras in Natural EnvironmentsabstractA new generation of inexpensive robotic pan-tilt cameras can maintain high-resolution panoramic displays of natural environments. However, the pan-tilt mechanisms are imprecise: small errors can produce large errors in the panoramic display. It is thus important to accurately estimate pan-tilt values. We present a new calibration algorithm that does not rely on calibration markers or fixed orthogonal edges which are rarely available in natural scenes. Our calibration algorithm uses image variance density to optimally estimate camera pan and tilt values by incrementally refining image registration using overlapping images from prior frames. Experiments suggest that the new calibration algorithm can reduce calibration error by 81%. In a companion paper, we present a new image registration algorithm based on spherical projection that optimally aligns the resulting frames Dezhen Song, Ni Qin, Kenneth Y. Goldberg |
ICRA | 1 |
| 2006 | Trajectory Tracking and Balance Stabilization Control of Autonomous MotorcyclesabstractWe report a new trajectory tracking and balancing control algorithm for an autonomous motorcycle. Building on the existing modeling work of a bicycle, the new dynamic model of the autonomous motorcycle considers the bicycle caster angle and captures the steering effect on the vehicle tracking and balancing. The trajectory tracking control takes an external/internal model decomposition approach. A nonlinear controller is designed to handle the vehicle balancing. The motorcycle balancing is guaranteed by the system internal equilibria calculation and by the trajectory and system dynamics requirements. The proposed control system is validated by numerical simulations, and is based on a real prototype motorcycle system Jingang Yi, Dezhen Song, Anthony Levandowski, Suhada Jayasuriya |
ICRA | 2 |
| 2006 | Vision-based Motion Planning for an Autonomous Motorcycle on Ill-Structured RoadabstractWe report our development of a vision-based motion planning system for an autonomous motorcycle designed for desert terrain, where uniform road surface and lane markings are not present. The motion planning is based on a vision vector space (V2-Space), which is an unitary vector set that represents local collision-free directions in the image coordinate system, V2-Space is constructed by extracting the vectors based on the similarity of adjacent pixels, which captures both the color information and the directional information from prior vehicle tire tracks and pedestrian footsteps. We report how V2-Space is constructed to reduce the impact of varying lighting conditions in outdoor environments. We also show how V2-Space can be used to incorporate vehicle kinematic, dynamic, and time-delay constraints in motion planning to fit the highly dynamic requirements of the motorcycle. The combined algorithm of the V2-Space construction and the motion planning runs in O(n) time, where n is the number of pixels in the captured image. Experiments show that our algorithm outputs correct robot motion commands more than 90% of the time Dezhen Song, Hyun Nam Lee, Jingang Yi, Anthony Levandowski |
IROS | 1 |
| 2006 | Exact algorithms for single frame selection on multiaxis SatellitesabstractNew multi-axis satellites allow camera imaging parameters to be set during each time slot based on competing demand for images, specified as rectangular requested viewing zones over the camera's reachable field of view. The single frame selection (SFS) problem is to find the camera frame parameters that maximize reward during each time window. We formalize the SFS problem based on a new reward metric that takes into account area coverage and image resolution. For a set of n client requests and a satellite with m discrete resolution levels, we give an algorithm that solves the SFS problem in time O(n/sup 2/m). For satellites with continuously variable resolution (m=/spl infin/), we give an algorithm that runs in time O(n/sup 3/). We have implemented all algorithms and verify performance using random inputs. Note to Practitioners-This paper is motivated by recent innovations in earth imaging by commercial satellites. In contrast to previous methods that required waits of up to 21 days for desired earth- satellite alignment, new satellites have onboard pan-tilt-zoom cameras that can be remotely directed to provide near real-time response to requests for images of specific areas on the earth's surface. We consider the problem of resolving competing requests for images: Given client demand as a set of rectangles on the earth surface, compute camera settings that optimize the tradeoff between pan, tilt, and zoom parameters to maximize camera revenue during each time slot. We define a new quality metric and algorithms for solving the problem for the cases of discrete and continuous zoom values. These results are a step toward multiple frame selection which will be addressed in future research. The metric and algorithms presented in this paper may also be applied to collaborative teleoperation of ground-based robot cameras for inspection and videoconferencing and for scheduling astronomic telescopes. Dezhen Song, A. Frank van der Stappen, Kenneth Y. Goldberg |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2005 | Steady-State Throughput and Scheduling Analysis of Multi-Cluster Tools for Semiconductor Manufacturing: A Decomposition ApproachabstractCluster tools are widely used as semiconductor manufacturing equipment. While throughput analysis and scheduling of single-cluster tools have been well-studied, the corresponding research on multi-cluster tools is still at the early stage. This paper analyzes steady-state throughput and scheduling of multi-cluster tools. A decomposition method is utilized to reduce a multi-cluster tool problem to multiple single-cluster tool problems. Existing research on the throughput and scheduling results is then applied to each single-cluster tool. For an M-cluster tool, an O(M) throughput calculation and robot scheduling algorithm is presented. A chemical-mechanical planarization (CMP) polisher is used as an example of the multi-cluster cluster tools to illustrate the proposed decomposition method and algorithms. Jingang Yi, Shengwei Ding, Dezhen Song |
ICRA | 3 |
| 2005 | Networked Robotic Cameras for Collaborative Observation of Natural Environments
Dezhen Song, Kenneth Y. Goldberg |
ISRR | 1 |
| 2004 | Unsupervised Scoring for Scalable Internet-based Collaborative TeleoperationabstractFor applications in education and entertainment, scalable Internet-based collaborative teleoperation allows many users simultaneously to share control of a single device. Automated numerical methods that can assess and record performance provide an incentive for users to participate and a means to evaluate individual and group performance. In this paper we describe "unsupervised scoring": a numerical approach to assessment based on clustering and response time. Like unsupervised learning, this approach is based on identifying regularities in the input rather than comparing input with desired output specified by an external supervisor. We present an algorithm for rapidly computing user scores that scales linearly with the number of users. We describe an implemented Java-based user interface incorporating this metric, an application based on the classic Twister game, and results where individual scores are compared with group performance. Kenneth Y. Goldberg, Dezhen Song, In Yong Song, Jane McGonigal, Dana Plautz |
ICRA | 2 |
| 2004 | An Exact Algorithm Optimizing Coverage-resolution for Automated Satellite Frame SelectionabstractNear real time satellite imaging provides timely images of the earth for weather prediction, disaster response, search and rescue, surveillance, and defense applications. As the satellite passes over the earth, camera imaging parameters are changed during each time window based on demand for images, specified as user requested zones in the reachable field of view during that time window. The satellite frame selection (SFS) problem is to find the camera frame parameters that maximize reward during each time window. To automate satellite management, we formalize the SFS problem based on a new reward metric that incorporates both image resolution and coverage. For a set of n client requests we give a series of algorithms, the fastest computes optimal results in O(n/sup 3/) for satellites with continuously variable resolution. We have implemented the algorithms and compare computation speed for all algorithms. Dezhen Song, A. Frank van der Stappen, Kenneth Y. Goldberg |
ICRA | 1 |
| 2003 | Efficient algorithms for shared camera controlabstractWe consider a system that allows n networked users to share control over a robotic webcamera. Each user guides the camera pan, tilt and zoom, by drawing a rectangle in the user interface. The server adjusts the camera to best satisfy the user requests, by solving a geometric optimization problem that requires fitting one rectangle to many. We improve upon previous results with an O(n3/2 log3 n) time exact algorithm for this problem. We also present a simple near-linear time e-approximation algorithm. We have implemented the latter and report on experimental results. Sariel Har-Peled, Vladlen Koltun, Dezhen Song, Kenneth Y. Goldberg |
SCG | 3 |
| 2003 | ShareCam part 1: interface, system architecture, and implementation of a collaboratively controlled robotic WebcamabstractShareCam is a robotic pan, tilt, and zoom web-based camera controlled by simultaneous frame requests from online users. Part II describes algorithms. This paper, part I, focuses on the system. Robotic Webcameras are commercially available but currently restrict control only one user at a time. ShareCam introduces a new interface that allows simultaneous control many users. In this Java-based interface, participating users interact desired frames remotely located browsers where users draw desired frames over a fixed panoramic image. User inputs re transmitted back to pair of PC servers that compute optimal camera back to a pair of PC servers that compute optimal camera parameters, servo the camera, and provide a video stream to all users. We describe the system, online experiments, and compare results with two frame selection models based on user "satisfaction", one memoryless and the second based on satisfaction over multiple motion cycles. Dezhen Song, Kenneth Y. Goldberg |
IROS | 1 |
| 2003 | ShareCam part II: approximate and distributed algorithms for a collaboratively controlled robotic WebcamabstractShareCam is a robotic pan, tilt, and zoom Web-based camera controlled by simultaneous frame requests from online users. Part I describes the system. This paper, part II, focuses on algorithms. The ShareCam problem is to find a camera frame that optimizes a measure of total user satisfaction. We present a grid-based approximation algorithm: given camera frame requests from n users, and approximation bound /spl epsi/, we analyze the trade of between solution quality and processing speed and prove that the algorithm runs in O(n//spl epsi//sup 3/) time. The algorithm can be distributed to run in O(1//spl epsi//sup 3/) time at each client and in O(n + 1//spl epsi//sup 3/) time at the server. Experiments suggest that performance of the distributed algorithm degrades gracefully as clients fail to complete their part of the computation. Dezhen Song, Anatoly Pashkevich, Kenneth Y. Goldberg |
IROS | 1 |
| 2003 | Algorithms and systems for shared access to a robotic streaming video cameraabstractINTRODUCTION Robotic streaming video cameras with pan, tilt, and zoom controls are now commercially available and are being installed in hundreds of locations around the world . Remote viewers can adjust camera parameters via the Internet to observe desired details in the scene. Current methods restrict control to one user at a time; users have to wait in a queue for their turn to operate the camera. In this thesis, we develop ShareCam, a new approach that eliminates the queue and allows many users to access and share control of the robotic camera simultaneously. Since conflicting frame requests are made by users, a primary challenge is computing optimal camera parameters. We formalize the problem using a new metric, Intersection Over Maximum (IOM), to model the degree of satisfaction for each user, and seek to maximize total satisfaction for n users. We develop online algorithms to solve this optimization problem for cases where pan, tilt and zoom values can be either discrete or Dezhen Song |
ACM Multimedia | 1 |
| 2003 | The co-opticon: shared access to a robotic streaming video cameraabstractThe "co-opticon" is a robotic pan, tilt, and zoom streaming video camera controlled by simultaneous frame requests from remote users. Robotic webcameras are commercially available but currently restrict control to only one user at a time. The co-opticon introduces a new interface that allows simultaneous control by many users. We will demonstrate the implemented system using a Java-based interface at the conference linked via the Internet to a camera on the UC Berkeley campus. We will also discuss system architecture and several new algorithms we've developed to compute optimal camera paramters based on user frame requests. The co-opticon can be tested online at: www.tele-actor.net/co-opticon. Dezhen Song, Kenneth Y. Goldberg |
ACM Multimedia | 1 |
| 2003 | Collaborative teleoperation using networked spatial dynamic votingabstractWe describe a networked teleoperation system that allows groups of participants to collaboratively explore live remote environments. Participants collaborate using a spatial dynamic voting (SDV) interface that allows them to vote on a sequence of images via a network such as the Internet. The SDV interface runs on each client computer and communicates with a central server that collects, displays, and analyzes time sequences of spatial votes. The results are conveyed to the "tele-actor", a skilled human with cameras and microphones who navigates and performs actions in the remote environment. This paper formulates analysis in terms of spatial interest functions and consensus regions, and presents system architecture, interface, and algorithms for processing voting data. Kenneth Y. Goldberg, Dezhen Song, Anthony Levandowski |
Proc. IEEE | 2 |
| 2002 | Collaborative Online Teleoperation with Spatial Dynamic Voting and a Human "Tele-Actor"abstractInternet-based "online robots" now provide public access to remote locations such as museums and laboratories. The Tele-Actor is a collaborative online teleoperation system for distance learning that allows many students to simultaneously share control of a single mobile resource. Our goal is to preserve the educational advantages of field trips without the drawbacks of group travel. We propose the "spatial dynamic voting" (SDV) interface for multiple operator single robot (MOSR) teleoperation. The SDV collects, displays, and analyzes a sequence of spatial votes from multiple online operators at their Internet browsers. The votes drive the motion of a single mobile robot or human "Tele-Actor". The paper describes Version 3.0 of the system architecture, SDV interface, algorithms for automated goal selection, and metrics for collaboration and leadership. We report results from a July 2001 field test with 56 remote users. Kenneth Y. Goldberg, Dezhen Song, Yoek-Nam Khor, David Pescovitz, Anthony Levandowski, Jesse C. Himmelstein, Janice Shih, Annamarie Ho, Eric Paulos, Judith S. Donath |
ICRA | 2 |
| 2002 | Exact and Distributed Algorithms for Collaborative Camera Control
Dezhen Song, A. Frank van der Stappen, Kenneth Y. Goldberg |
WAFR | 1 |