Di Wang 0020

dblp:18/5410-20 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-6998-7718ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Systems, architecture and hardware · 14 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Heterogeneous Sensor Fusion and Active Perception for Transparent Object Reconstruction with a PDM2 Sensor and a Camera
abstract
Transparent household objects present a challenge for domestic service robots, since neither regular cameras nor RGB-D cameras can provide accurate points for shape reconstruction. The new type of pretouch dual-modality distance and material sensor (PDM2) can provide reliable and accurate depth readings, but it is a point sensor and scanning the object exclusively with the sensor is too inefficient. Hence, we present a sensor fusion approach by combining a regular camera with the PDM2sensor. The approach is based on a data fusion algorithm for shape reconstruction and an active perception algorithm for scan planning for the PDM2sensor. The data fusion algorithm is a distributed Gaussian process (GP)-based shape reconstruction method that allows for incremental local update to reduce computational time. The active perception algorithm is an optimization-based approach by increasing the information gain (IG) and prioritizing the boundary points under a preset travel distance constraint. We have implemented and tested the algorithms with six different transparent household items. The results show satisfactory shape reconstruction results in all test cases with an average increase in intersection over union (IoU) from 0.73 to 0.96.
Fengzhi Guo, Shuangyu Xie, Di Wang 0020, Dezhen Song
ICRA3
2024 Coupled Active Perception and Manipulation Planning for a Mobile Manipulator in Precision Agriculture Applications
abstract
A mobile manipulator often finds itself in an application where it needs to take a close-up view before performing a manipulation task. Named this as a coupled active perception and manipulation (CAPM) problem, we model the uncertainty in the perception process and devise a key state/task planning algorithm that considers reachability conditions jointly established from perception and manipulation task constraints. By minimizing expected energy usage in body key state planning while satisfying task constraints, our algorithm is able to find an energy-efficient trajectory with less body repositioning motion while ensuring the success of the task. We have implemented the algorithm and tested it in both simulation and physical experiments. The results have confirmed that our algorithm has a lower energy consumption compared to a two-stage decoupled approach, while still maintaining a success rate of 100% for the task.
Shuangyu Xie, Chengsong Hu, Di Wang 0020, Joe Johnson, Muthukumar Bagavathiannan, Dezhen Song
ICRA3
2024 Toward Precise Robotic Weed Flaming Using a Mobile Manipulator with a Blowtorch
abstract
Robotic weed flaming is a new and environmentally friendly approach to weed removal in the agricultural field. Using a mobile manipulator equipped with a blowtorch, we design a new system and algorithm to enable effective weed flaming, which requires robotic manipulation with a soft and deformable end effector, as the thermal coverage of the flame is affected by dynamic or unknown environmental factors such as gravity, wind, atmospheric pressure, fuel tank pressure, and pose of the nozzle. System development includes overall design, hardware integration, and software pipeline. To enable precise weed removal, the greatest challenge is to detect and predict dynamic flame coverage in real time before motion planning, which is quite different from a conventional rigid gripper in grasping or a spray gun in painting. Based on the images from two onboard infrared cameras and the pose information of the blowtorch nozzle on a mobile manipulator, we propose a new dynamic flame coverage model. The flame model uses a center-arc curve with a Gaussian cross-section model to describe the flame coverage in real time. The experiments have demonstrated the working system and shown that our model and algorithm can achieve a mean average precision (mAP) of more than 76% in the reprojected images during online prediction.
Di Wang 0020, Chengsong Hu, Shuangyu Xie, Joe Johnson, Hojun Ji, Yingtao Jiang, Muthukumar Bagavathiannan, Dezhen Song
IROS1
2023 The Third Generation (G3) Dual-Modal and Dual Sensing Mechanisms (DMDSM) Pretouch Sensor for Robotic Grasping
abstract
Fingertip-mounted pretouch sensors are very useful for robotic grasping. In this paper, we report a new (G3) dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor for near-distance ranging and material sensing, which is based on pulse-echo ultrasound (US) and optoacoustics (OA). Different from previously reported versions, the G3 sensor utilizes a self-focused US/OA transceiver, thereby eliminating the need of a bulky parabolic reflective mirror for focusing the ultrasound and laser beams. The self-focused laser and ultrasound beams can be easily steered by a (flat) scanning mirror which expands from single-point ranging and detection to areal mapping or imaging. To verify the new design, a prototype G3 DMDSM sensor with a scanning mirror is fabricated. The US and OA ranging performances are tested in experiments. Together with the scanning mirror, thin wire targets made of same or different materials at different positions are scanned and imaged. The ranging and imaging results show that the G3 DMDSM sensor can provide new and better pretouch mapping and imaging capabilities for robotic grasping than its predecessors.
Shuangliang Li, Di Wang 0020, Fengzhi Guo, Dezhen Song
ICRA3
2023 A Pretouch Perception Algorithm for Object Material and Structure Mapping to Assist Grasp and Manipulation Using a DMDSM Sensor
abstract
We report a new material and structure mapping (MSM) algorithm to assist robotic grasping and manipulation. Building on our new sensor development, the algorithm has four main components: 1) detection of time-of-flight (ToF) durations for the dual modalities of optoacoustic (OA) and pulse-echo ultrasound (US), 2) contour reconstruction by fusing OA and US signals, 3) local noise filtering by checking local consistency of material and structure label (MSL), and 4) medium boundary searching that identifies class boundaries through two-staged clustering and boundary establishment using support vector machine (SVM) hyperplanes. We have implemented our algorithm and tested it with multiple common household items. The experimental results have successfully validated our algorithm design which shows that the average error of contour reconstruction is 0.05 mm and the true positive rate of MSL is over 98%.
Fengzhi Guo, Shuangyu Xie, Di Wang 0020, Dezhen Song
IROS3
2022 The Second Generation (G2) Fingertip Sensor for Near-Distance Ranging and Material Sensing in Robotic Grasping
abstract
To continuously improve robotic grasping, we are interested in developing a contactless fingertip-mounted sensor for near-distance ranging and material sensing. Previously, we demonstrated a dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor prototype based on pulse-echo ultrasound and optoacoustics. However, the complex system, the bulky and expensive pulser-receiver, and the omni-directionally sensitive microphone block the sensor from practical applications in real robotic fingers. To address these issues, we report the second generation (G2) DMDSM sensor without the pulser-receiver and microphone, which is made possible by redesigning the ultrasound transmitter and receiver to gain much wider acoustic bandwidth. To verify our design, a prototype of the G2 DMDSM sensor has been fabricated and tested. The testing results show that the G2 DMDSM sensor can achieve better ranging and similar material/structure sensing performance, but with much-simplified configuration and operation. The primary results indicate that the G2 DMDSM sensor could provide a promising solution for fingertip pretouch sensing in robotic grasping.
Di Wang 0020, Dezhen Song
ICRA2
2021 Device Design and System Integration of a Two-Axis Water-immersible Micro Scanning Mirror (WIMSM) to Enable Dual-modal Optical and Acoustic Communication and Ranging for Underwater Vehicles
abstract
To address the communication and ranging challenges caused by underwater environment, we design dual modal devices for autonomous underwater vehicles (AUVs). The dual-modal design builds upon a co-axial ultrasonic and green laser beams which leverage different signal diverging patterns and different responses in the underwater environment by each modality to achieve robust adaptability. Here we report our recent progress in improving scanning and aiming capabilities for dual-modal beam steering. The core part is our Two-Axis Water-immersible Micro Scanning Mirror (WIMSM). We improve hinge design of WIMSM for larger scanning range. We incorporate high speed Hall effect sensor-based pose feedback channel to enable closed-loop scanning and aiming control. We design ultrasonic-assisted laser handshaking method to help AUVs to acquire optical underwater communication. We have prototyped our devices and tested them in a water tank. The initial results are promising.
Xiaoyu Duan, Di Wang 0020, Dezhen Song
ICRA2
2021 Fingertip Pulse-Echo Ultrasound and Optoacoustic Dual-Modal and Dual Sensing Mechanisms Near-Distance Sensor for Ranging and Material Sensing in Robotic Grasping
abstract
To improve robotic grasping, we are interested in developing a new non-contact fingertip-mounted sensor for near-distance ranging and material sensing. Here we report new progress in combining direct pulse-echo ultrasound and optoacoustic effects in sensor design to deal with optically and/or acoustically challenging targets (OACTs). Our dual-modal and dual sensing mechanisms (DMDSM) sensor design is enabled by a novel wideband ultrasound transmitter embedded inside a piezoelectric (lead zirconate titanate - PZT) ring transducer. The new DMDSM sensor is capable of differentiating a variety of OACTs. To verify our design, both distance ranging tests and material sensing tests have been conducted. The ranging tests show the sensor can perform both optoacoustic ranging (for light-absorbing materials) and pulse-echo ultrasound ranging (for reflective or transparent materials). For material sensing, the dual-modal spectra from OACTs are collected to compare the new sensor with previous designs. The overall 100% accuracy from the confusion matrices indicates the initial success of our sensor design in differentiating conventional targets as well as the OACTs with the new DMDSM sensor.
Di Wang 0020, Dezhen Song
ICRA2
2020 Fingertip Non-Contact Optoacoustic Sensor for Near-Distance Ranging and Thickness Differentiation for Robotic Grasping*
abstract
We report the feasibility study of a new optoacoustic sensor for both near-distance ranging and material thickness classification for robotic grasping. It is based on the optoacoustic effect where focused laser pulses are used to generate wideband ultrasound signals in the target. With a much smaller optical focal spot, the optoacoustic sensor achieves a lateral resolution of 93 μm, which is six times higher than ultrasound pulse-echo ranging under the same condition. A new multi-mode wideband PZT (lead zirconate titanate) transducer is built to properly receive the wideband optoacoustic signal. The ability to receive both low- and high-frequency components of the optoacoustic signal enhances the material sensing capability, which makes it promising to determine not only material type but also the sub-surface structures. For demonstration, optoacoustic spectra are collected from hard and soft materials with different thickness. A Bag-of-SFA-Symbols (BOSS) classifier is designed to perform primary material and then thickness classification based on the optoacoustic spectra. The accuracy of material / thickness classification reaches ≥ 99% and ≥ 94%, respectively, which shows the feasibility of differentiating solid materials with different thickness by the optoacoustic sensor.
Di Wang 0020, Dezhen Song
IROS2
2020 Toward Automatic Subsurface Pipeline Mapping by Fusing a Ground-Penetrating Radar and a Camera
abstract
We propose a novel subsurface pipeline mapping and 3D reconstruction method by fusing ground-penetrating radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with nonperpendicular angles. Furthermore, we fuse visual simultaneous localization and mapping outputs, encoder readings with GPR scans to classify hyperbolas into different pipeline groups. We extensively apply the J-linkage method and maximum likelihood estimation with error analysis to improve algorithm robustness and accuracy. As a result, we optimally estimate the radii and locations of all pipelines. We have implemented our method and tested it in physical experiments with representative pipeline configurations. Two different kinds of 3-m-long pipes are used, with radii being 4.62 and 3.02 cm, respectively. The results show that our method successfully reconstructs all subsurface pipes. Moreover, the average estimation errors for two orientation angles of pipelines are 1.73° and 0.73°, respectively. The average localization error is 4.47 cm. Note to Practitioners-Automatic and accurate underground pipeline mapping technology is very important in civil construction projects. Lack of 3D utility pipeline maps may lead to accidental damage in civil construction and maintenance. Although ground-penetrating radar (GPR-based pipeline mapping methods have been studied for several years, these methods require the perpendicular scanning with respect to the pipe, which is impossible to guarantee in practice since the orientations of pipelines are unknown. Furthermore, these traditional methods can only estimate one pipeline at a time in a survey area and require prior knowledge of pipe diameter. We propose a robotic subsurface pipeline mapping method with a GPR and a camera to handle difficult factors such as multiple pipes, unknown pipeline orientation, and unknown pipeline diameters. Hence, we can perform GPR scanning along any generic linear trajectories. Our method has been tested in physical experiments with representative pipeline configurations. The results are sufficiently accurate, and it proves that our method can be an effective technology to reconstruct the underground pipelines.
Haifeng Li 0008, Chieh Chou, Longfei Fan, Binbin Li 0006, Di Wang 0020, Dezhen Song
IEEE Trans Autom. Sci. Eng.5
2019 Toward Fingertip Non-Contact Material Recognition and Near-Distance Ranging for Robotic Grasping
abstract
We report the feasibility study of a new acoustic and optical bi-modal distance & material sensor for robotic grasping. The new sensor is designed to be mounted on the robot fingertip to provide last-moment perception before contact happens. It is based on both pulse-echo ultrasound and optoacoustic effects enabled by single-element air-coupled transducers. In contrast to conventional contact-based and recent pre-touch approaches, this new method overcomes their disadvantages and provides robotic fingers with the capability to detect the distance and material type of the target at a near distance before contact occurs, which is crucial for robust and nimble grasping. The proposed sensor has been tested with different materials, shapes, and porous properties. The experimental results show that this sensor design is functional and practical.
Di Wang 0020, Dezhen Song
ICRA2
2019 On the Tunable Sparse Graph Solver for Pose Graph Optimization in Visual SLAM Problems
abstract
We report a tunable sparse optimization solver that can trade a slight decrease in accuracy for significant speed improvement in pose graph optimization in visual simultaneous localization and mapping (vSLAM). The solver is designed for devices with significant computation and power constraints such as mobile phones or tablets. Two approaches have been combined in our design. The first is a graph pruning strategy by exploiting objective function structure to reduce the optimization problem size which further sparsifies the optimization problem. The second step is to accelerate each optimization iteration in solving increments for the gradient-based search in Gauss-Newton type optimization solver. We apply a modified Cholesky factorization and reuse the decomposition result from last iteration by using Cholesky update/downdate to accelerate the computation. We have implemented our solver and tested it with open source data. The experimental results show that our solver can be twice as fast as the counterpart while maintaining a loss of less than 5% in accuracy.
Chieh Chou, Di Wang 0020, Dezhen Song, Timothy A. Davis 0001
IROS2
2019 Virtual Lane Boundary Generation for Human-Compatible Autonomous Driving: A Tight Coupling between Perception and Planning
abstract
Existing autonomous vehicle (AV) navigation algorithms treat lane recognition, obstacle avoidance, local path planning, and lane following as separate functional modules which result in driving behavior that is incompatible with human drivers. It is imperative to design human-compatible navigation algorithms to ensure transportation safety. We develop a new tightly-coupled perception-planning framework that combines all these functionalities to ensure human-compatibility. Using GPS-camera-lidar sensor fusion, we detect actual lane boundaries (ALBs) and propose availability-reasonability-feasibility (ARF) threefold tests to determine if we should generate virtual lane boundaries (VLBs) or follow ALBs. If needed, VLBs are generated using a dynamically adjustable multi-objective optimization framework that considers obstacle avoidance, trajectory smoothness (to satisfy vehicle kinodynamic constraints), trajectory continuity (to avoid sudden movements), GPS following quality (to execute global plan), and lane following or partial direction following (to meeting human expectation). Consequently, vehicle motion is more human compatible than existing approaches. We have implemented our algorithm and tested under open source data with satisfying results.
Binbin Li 0006, Dezhen Song, Ankit Ramchandani, Hsin-Min Cheng, Di Wang 0020, Yiliang Xu, Baifan Chen
IROS5
2018 Encoder-Camera-Ground Penetrating Radar Tri-Sensor Mapping for Surface and Subsurface Transportation Infrastructure Inspection
abstract
We report system and algorithmic development for a sensing suite comprising multiple sensors for both surface and subsurface transportation infrastructure inspection focusing on multi-modal mapping for inspection. The sensing suite contains a camera, a ground penetrating radar (GPR), and a wheel encoder. We design the sensing suite and propose a data collection scheme using customized artificial landmarks (ALs). We use ALs to synchronize two types data streams: camera images that are temporally evenly-spaced and GPR/encoder data that are spatially evenly-spaced. We also employ pose graph optimization with synchronization as penalty functions to further refine synchronization and perform data fusion for 3D reconstruction. We have implemented the system and tested it in physical experiments. The results show that our system successfully fuses three sensory data and product metric 3D reconstruction. The sensor fusion approach reduces the end-to-end distance error from 7.45cm to 3.10cm.
Chieh Chou, Aaron Kingery, Di Wang 0020, Haifeng Li 0008, Dezhen Song
ICRA3
2018 Robotic Subsurface Pipeline Mapping with a Ground-penetrating Radar and a Camera
abstract
We propose a novel subsurface pipeline mapping method by fusing Ground Penetrating Radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with non-perpendicular angles. Furthermore, we fuse visual simultaneous localization and mapping outputs, encoder readings with GPR scans to classify hyperbolas into different pipeline groups. We extensively apply the J-Linkage method and maximum likelihood estimation to improve algorithm robustness and accuracy. As the result, we optimally estimate the radii and locations of all pipelines. We have implemented our method and tested it in physical experiments with representative pipeline configurations. The results show that our method successfully reconstructs all subsurface pipes. Moreover, the average localization error is 4.69cm.
Haifeng Li 0008, Chieh Chou, Longfei Fan, Binbin Li 0006, Di Wang 0020, Dezhen Song
IROS5