Chieh-Chih Wang

dblp:53/4755 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-9385-0044ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 4 first-author · 10 since 2021Systems, architecture and hardware · 27 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 TS-DETR: Traffic Sign Detection Based on Positive and Negative Sample Augmentation
abstract
Traffic sign detection plays an essential role in advanced driver assistance system (ADAS) or self-driving vehicles. Typically, deep neural networks are employed to analyze road scene images captured by an onboard camera. However, due to the significant variation in appearance of different traffic signs, the classification of high similarity patterns is still a challenging task. To address these issues, this paper presents an end-to-end traffic sign detection framework based on DETR. The proposed network incorporates data augmentation and negative sample learning to mitigate the problem of data imbalance and enhance the model recognition capability effectively. An UASPP module (Upsample Atrous Pyramid Pooling) is introduced to integrate multi-scale features and global information. In the experiments, the performance evaluation has demonstrated the improvement of mAP by 3.9% on TT100K and 36.3% on GTSDB compared to state-of-the-art methods. The code and datasets are available at https://github.com/chinglun/TS-DETR.
Ching-Lun Lin, Huei-Yung Lin, Chieh-Chih Wang
ICRA3
2025 FuseRoad: Enhancing Lane Shape Prediction Through Semantic Knowledge Integration and Cross-Dataset Training
abstract
The rapid evolution of advanced driver assistance systems (ADAS) has been driven by the advances of deep neural networks, and multi-tasking is essential for autonomous driving systems. This paper presents FuseRoad, a new multi-task model that leverages cross-dataset learning to address the dependency on specific multi-task datasets and reduce the annotation costs. It integrates semantic segmentation and lane detection into an end-to-end framework while providing an effective approach to utilize multiple single-task datasets. By incorporating Semantic Road Knowledge Extractor (SRKE) to direct more attentions on the roadway, FuseRoad enhances the accuracy and reliability of lane detection. The model also employs the logit normalization loss to address the issue of overconfidence commonly faced by conventional lane detection methods. In experiments, FuseRoad outperforms state-of-the-art approaches in both accuracy and F-1 score. The evaluation on semantic segmentation metrics also demonstrates that the proposed technique is highly effective for multi-task road scene analysis. Code and datasets are available at https://github.com/HengChihHsiao/FuseRoad.
Heng-Chih Hsiao, Yi-Chang Cai, Huei-Yung Lin, Walon Wei-Chen Chiu, Chiao-Tung Chan, Chieh-Chih Wang
IV6
2024 Self-Supervised Motion Segmentation with Confidence-Aware Loss Functions for Handling Occluded Pixels and Uncertain Optical Flow Predictions
abstract
In driving scenarios, motion segmentation is a crucial and fundamental component that is needed for many tasks. Recently, a self-supervised multitasking framework was proposed for driving scenarios. It simultaneously trains motion segmentation, optical flow, depth, and ego-motion models without annotated data. The self-supervised architecture derives training signals from training data via loss functions. If these loss functions lack robustness, they may result in model inaccuracies. To reduce the bad influences of occlusion and optical flow estimation errors on motion segmentation, we propose two loss functions: (1) Soft-Per-Pixel-Minimum (Soft-PPM) loss that excludes occluded pixels while balancing the contribution of each frame on the loss function temporally; (2) Flow difference loss that excludes pixels with unclear motion states to diminish the effect of optical flow estimation errors. Our loss function design is based on the key insight that information such as depth and optical flow can be used to train motion segmentation models and act as a reliable measure for pixels during training. Our approach can improve segmentation accuracy for both moving and static objects and has achieved IoU scores on moving and static classes comparable to the state-of-the-art methods on the KITTI dataset.
Chung-Yu Chen, Bo-Yun Lai, Ying-Shiuan Huang, Wen-Chieh Lin, Chieh-Chih Wang
IROS5
2024 Multi-modal Motion Prediction using Temporal Ensembling with Learning-based Aggregation
abstract
Recent years have seen a shift towards learning-based methods for trajectory prediction, with challenges remaining in addressing uncertainty and capturing multi-modal distributions. This paper introduces Temporal Ensembling with Learning-based Aggregation, a meta-algorithm designed to mitigate the issue of missing behaviors in trajectory prediction, which leads to inconsistent predictions across consecutive frames. Unlike conventional model ensembling, temporal ensembling leverages predictions from nearby frames to enhance spatial coverage and prediction diversity. By confirming predictions from multiple frames, temporal ensembling compensates for occasional errors in individual frame predictions. Furthermore, trajectory-level aggregation, often utilized in model ensembling, is insufficient for temporal ensembling due to a lack of consideration of traffic context and its tendency to assign candidate trajectories with incorrect driving behaviors to final predictions. We further emphasize the necessity of learning-based aggregation by utilizing mode queries within a DETR-like architecture for our temporal ensembling, leveraging the characteristics of predictions from nearby frames. Our method, validated on the Argoverse 2 dataset, shows notable improvements: a 4% reduction in minADE, a 5% decrease in minFDE, and a 1.16% reduction in the miss rate compared to the strongest baseline, QCNet, highlighting its efficacy and potential in autonomous driving.
Kai-Yin Hong, Chieh-Chih Wang, Wen-Chieh Lin
IROS2
2024 Enhancing LiDAR Scene Upsampling with Instance-aware Feature-embedding and Attention Mechanism
abstract
Scanning LiDAR is one of the widely used sensors in autonomous vehicles; however, the inherent sparsity of LiDAR point clouds often affects its performance. To address this issue, upsampling methods could be employed to enhance low-resolution LiDAR data. Although there have been methods on upsampling of single-object point clouds recently in computer vision, they tend to generate a considerable amount of artifacts when dealing with real-world LiDAR scenes consisting of multiple objects. In this paper, we propose a solution to tackle this problem by introducing an instance embedding auxiliary task and a context attention module. With our auxiliary learning architecture, the network can learn features that benefit both the primary upsampling task and the auxiliary instance embedding task. This training design enables the point generation process to be carried out separately and significantly reduces artifacts of the upsampling results on the SemanticKITTI dataset, particularly in areas surrounding instances. By leveraging these techniques to improve the model’s understanding of the relationship between objects and the background in LiDAR scenes, we achieve an overall 4% to 10% improvement in whole-scene upsampling.
Wei-Jen Wang, You-Sheng Do, Wen-Chieh Lin, Chieh-Chih Wang
IROS4
2023 GNN-Based Point Cloud Maps Feature Extraction and Residual Feature Fusion for 3D Object Detection
abstract
LiDAR detection of long-range vehicles is challenging because very few and sparse points are measured in long distances and vehicles with similar shapes of targets could lead to false positives easily. To tackle these challenges, taking the environment information (HD maps) into account could be beneficial to predetermine where targets are more or less likely to appear. Compared with semantic maps, HD maps formed by point clouds provide much richer information from surrounding static objects and scenes. In this work, we construct a GNN-based feature extraction of point cloud maps to increase the receptive fields of learning map features. Our work is based on PVRCNN, the state-of-the-art LiDAR object detection method. With point-wise and voxel-wise features obtained from PVRCNN, residual feature fusion is proposed to fuse the features from PVRCNN and the map features from GNN. Our approach is evaluated on NuScenes dataset. It achieves a 24.78% average precision improvement for long-range objects at 40–50 meters, the farthest areas with ground truth annotation. Our approach also has a 4.22% reduction of false positives in the entire sensing areas.
Wei-Hsiang Liao 0005, Chieh-Chih Wang, Wen-Chieh Lin
ICRA2
2023 Asynchronous State Estimation of Simultaneous Ego-motion Estimation and Multiple Object Tracking for LiDAR-Inertial Odometry
abstract
We propose LiDAR-Inertial Odometry via Simultaneous EGo-motion estimation and Multiple Object Tracking (LIO-SEGMOT), an optimization-based odometry approach targeted for dynamic environments. LIO-SEGMOT is formulated as a state estimation approach with asynchronous state update of the odometry and the object tracking. That is, LIO-SEGMOT can provide continuous object tracking results while preserving the keyframe selection mechanism in the odometry system. Meanwhile, a hierarchical criterion is designed to properly couple odometry and object tracking, preventing system instability due to poor detections. We compare LIO-SEGMOT against the baseline model LIO-SAM, a state-of-the-art LIO approach, under dynamic environments of the KITTI raw dataset and the self-collected Hsinchu dataset. The former experiment shows that LIO-SEGMOT obtains an average improvement 1.61% and 5.41% of odometry accuracy in terms of absolute translational and rotational trajectory errors. The latter experiment also indicates that LIO-SEGMOT obtains an average improvement 6.97% and 4.21% of odometry accuracy.
Yu-Kai Lin, Wen-Chieh Lin, Chieh-Chih Wang
ICRA3
2023 Lidar-Based Multiple Object Tracking with Occlusion Handling
abstract
Occlusion remains an issue in multiple object tracking, which could cause ambiguity in object detection, such as incorrect or missing detection. Under occlusion, a track could experience an early termination, resulting in identity switches and/or fragmentation. To recover from different lengths of occlusions, the track should be maintained by considering its occlusion status. To address the issues mentioned above, we propose an indicator that can model the track's occlusion extent via geometric information provided by LiDAR data. Through incorporating the indicator into the track management and data association process, it is feasible to prevent tracks from premature termination. The proposed method is evaluated on the collected dataset which undergoes frequent and severe occlusions. Compared to the state-of-the-art probabilistic tracking approach, our approach achieves improvements of 3.26% in MOTA and 5.36% in IDF1. Additionally, we obtain 9.89% improvements in IDF1 specifically for objects experiencing severe occlusions.
Ruo-Tsz Ho, Chieh-Chih Wang, Wen-Chieh Lin
IROS2
2023 Automotive Radar Missing Dimension Reconstruction from Motion
abstract
Automotive radars have been reliably used in most assisted and autonomous driving systems due to their robustness to extreme weather conditions. With radial velocity measurements from automotive radars, moving targets such as cars, trucks, and buses can be tracked robustly. However, due to the lack of elevation angles in measurements from automotive radars, stationary targets at different heights, such as maintenance holes and bridges, cannot be distinguished. Most autonomous systems rely on sensor fusion or ignore stationary targets to tackle the problem of missing the elevation angle dimension, which derives safety issues. We propose a simple yet effective approach to estimate the elevation angles of stationary targets from relative velocity and radial velocity measurements from an automotive radar. In contrast to structure from motion in computer vision, we utilize the instantaneous velocity generated from the motion of the ego vehicle. The radial velocity of each target is the projection of relative velocity onto the radial direction from radar to target. The radial velocity of each target can be inferred given the target's azimuth, elevation angle, and relative velocity. Accordingly, the elevation angle of each stationary target can be uniquely calculated given the velocity of radar and the target's azimuth and radial velocity measurements. The radar's velocity is estimated with the existing radar odometry algorithm and IMU. The proposed method is verified with real-world data. We evaluate the system's performance with a pre-built point cloud map and a good localization module in a real-world scenario. The proposed elevation angle reconstruction can reach a 1.41-degree mean error and standard deviation of 0.6 degrees in elevation angle.
Chun-Yu Hou, Chieh-Chih Wang, Wen-Chieh Lin
IROS2
2021 A Normal Distribution Transform-Based Radar Odometry Designed For Scanning and Automotive Radars
abstract
Existing radar sensors can be classified into automotive and scanning radars. While most radar odometry (RO) methods are only designed for a specific type of radar, our RO method adapts to both scanning and automotive radars. Our RO is simple yet effective, where the pipeline consists of thresholding, probabilistic submap building, and an Normal Distribution Transform-based (NDT-based) radar scan matching. The proposed RO has been tested on two public radar datasets: the Oxford Radar RobotCar dataset and the nuScenes dataset, which provide scanning and automotive radar data respectively. The results show that our approach surpasses state-of-the-art RO using either automotive or scanning radar by reducing translational error by 51% and 30%, respectively, and rotational error by 17% and 29%, respectively. Besides, we show that our RO achieves centimeter-level accuracy as lidar odometry, and automotive and scanning RO have similar accuracy.
Pou-Chun Kung, Chieh-Chih Wang, Wen-Chieh Lin
ICRA2
2020 Extrinsic and Temporal Calibration of Automotive Radar and 3D LiDAR
abstract
While automotive radars are widely used in most assisted and autonomous driving systems, only a few works were proposed to tackle the calibration problems of automotive radars with other perception sensors. One of the key calibration challenges of automotive planar radars with other sensors is the missing elevation angle in 3D space. In this paper, extrinsic calibration is accomplished based on the observation that the radar cross section (RCS) measurements have different value distributions across radar's vertical field of view. An approach to accurately and efficiently estimate the time delay between radars and LiDARs based on spatial-temporal relationships of calibration target positions is proposed. In addition, a localization method for calibration target detection and localization in pre-built maps is proposed to tackle insufficient LiDAR measurements on calibration targets. The experimental results show the feasibility and effectiveness of the proposed Radar-LiDAR extrinsic and temporal calibration approaches.
Chia-Le Lee, Yu-Han Hsueh, Chieh-Chih Wang, Wen-Chieh Lin
IROS3
2017 Cross-Device Wi-Fi Map Fusion with Gaussian Processes
abstract
spatially sparse received signal strength measurements obtained with multiple devices. First, we show that the residual of the linear regression between devices, usually unaccounted for in existing cross-device localization work, is an important indicator of device dissimilarity and a good predictor of localization performance. Through explicitly modeling the device dissimilarity, one can improve localization accuracy when fusing training sets from multiple devices by weighting each training set differently. Second, we use the Gaussian process (GP) sensor model to develop a regression algorithm which more reliably estimates the linear fit and device dissimilarity given only a few labeled samples from each new device. By accounting for device dissimilarities in map fusion and by using the proposed regression algorithm, localization performance can be greatly improved given just a few training samples from a new device. Also, when fusing multiple existing maps for a new device using regression misfit, performance is improved by 3.5 to 10 percent.
Hsiao-Chieh Yen, Chieh-Chih Wang
IEEE Trans. Mob. Comput.2
2016 Exploiting Moving Objects: Multi-Robot Simultaneous Localization and Tracking
abstract
Cooperative localization has been proved to effectively outperform single-robot localization. While most of the state-of-the-art multi-robot localization systems either treat moving objects as outliers or accomplish moving object tracking separately from localization, we argue that augmenting moving objects into the localization estimation can further enhance localization performance and is indeed the key to solve several localization challenges such as insufficient map features, no map features, and symmetric maps. In this paper, a multi-robot simultaneous localization and tracking (MR-SLAT) algorithm based on the extended Kalman filter is proposed, and multiple hypothesis tracking (MHT) is integrated into MR-SLAT for dealing with challenging data association issues. The proposed approach is verified in two scenarios: the NAO humanoid robots equipped with cameras and WiFi are used in the RoboCup scenario and the robotic vehicles with laser scanners and dedicated short-range communications (DSRC) are used in the traffic scenario. The experiments with ground truth show that MR-SLAT, by exploiting moving objects, is superior to single-robot localization and cooperative localization in challenging scenarios. Ample experimental and simulation results demonstrate the effectiveness of exploiting moving objects and the generality and feasibility of the proposed MR-SLAT algorithm.
Chun-Hua Chang, Shao-Chen Wang, Chieh-Chih Wang
IEEE Trans Autom. Sci. Eng.3
2015 2-point RANSAC for scene image matching under large viewpoint changes
abstract
This work aims to accurately match two scene images under large viewpoint changes, which is the key issue in appearance-based localization tasks. In this paper, two key ideas are proposed to solve the challenging problem. First, to detect extreme small overlapping regions between two images, a new approach is developed to estimate the camera motion using only two pairs of matched features, while the state-of-art needs at least five. Second, proper prior knowledge to the environmental structures is utilized to strengthen the outlier rejection. The proposed 2-point approach is tested on challenging scenes and shows good robustness to the drastic occlusion and scaling caused by viewpoint changes.
Chih-Chung Chou, Chieh-Chih Wang
ICRA2
2014 Fisher's Discriminant with Natural Image Priors
abstract
Linear discriminant analysis that takes spatial smoothness into account has been developed and widely used in image processing society. However, two questions remain unanswered. First, which is the best way to incorporate the smoothness property of images with linear discriminant analysis? Second, which is the best representation for the smoothness property of images? To answer the first question, we propose a Bayesian framework of Gaussian process in order to extend Fisher's discriminant for image data. The probability structure for our extended Fisher's discriminant is explicitly formulated, and the smoothness properties of images are utilized as prior probabilities. For the second question, we suggest a family of prior probabilities derived from natural image statistics. The unknown parameters in our model are estimated via the maximum a posteriori probability (MAP) estimation. We will show that existing methods imposing smoothness assumption of images are rough approximations to the proposed MAP estimates in this framework. Experimental results on the Yale face database and the ETH-80 object categorization dataset show that the proposed method significantly outperforms the other Fisher's discriminant methods for various image data.
Yao-Hsiang Yang, Lu-Hung Chen, Chu-Song Chen, Chieh-Chih Wang
ICPR4
2014 Communication adaptive multi-robot simultaneous localization and tracking via hybrid measurement and belief sharing
abstract
Existing multi-robot cooperative perception solutions can be mainly classified into two categories, measurement-based and belief-based, according to the information shared among robots. With well-controlled communication, measurement-based approaches are expected to achieve theoretically optimal estimates while belief-based approaches are not because the cross-correlations between beliefs are hard to be perfectly estimated in practice. Nevertheless, belief-based approaches perform relatively stable under unstable communication as a belief contains the information of multiple previous measurements. Motivated by the observation that measurement sharing and belief sharing are respectively superior in different conditions, in this paper a hybrid algorithm, communication adaptive multi-robot simultaneous localization and tracking (ComAd MR-SLAT), is proposed to combine the advantages of both. To tackle the unknown or unstable communication conditions, the information to share is decided by maximizing the expected uncertainty reduction online, based on which the algorithm dynamically alternates between measurement-sharing and belief-sharing without information loss or reuse. The proposed ComAd MR-SLAT is evaluated in communication conditions with different packet loss rates and bursty loss lengths. In our experiments, ComAd MR-SLAT outperforms measurement-based and belief-based MR-SLAT in accuracy. The experimental results demonstrate the effectiveness of the proposed hybrid algorithm and exhibit that ComAd MR-SLAT is robust under different communication conditions.
Chun-Kai Chang, Chun-Hua Chang, Chieh-Chih Wang
ICRA3
2014 Deep learning of spatio-temporal features with geometric-based moving point detection for motion segmentation
abstract
This paper introduces an approach to accomplish motion segmentation from a moving stereo camera based on deep learning. Previous work on moving object detection mostly use point features based on 3D geometric constraints. However, point features require good features, and are hard to detect or to be matched correctly in situations where objects have smooth textures. To alleviate this problem, learning high-level spatio-temporal features unsupervisedly from raw image data based on Reconstruction Independent Component Analysis (RICA) autoencoders is proposed. Despite the power of the new spatio-temporal features, these features cannot not learn and be used to interpret 3D geometry of dynamic scenes, which is critical for moving object detection from moving cameras. As detected moving points based on 3D geometric constraints still contain valuable information of 3D scene as well as the camera egomotion, we propose a framework that incorporates both the detected moving point results and the learned spatio-temporal features as inputs to Recursive Neural Networks (RNN) that performs motion segmentation. Both features effectively complement each other. The proposed approach is demonstrated with real-world stereo video data that contains multiple moving objects, and has achieved 26% better detection rate over the existing 3D geometric-based moving points detector.
Tsung-Han Lin, Chieh-Chih Wang
ICRA2
2013 Adapting Gaussian processes for cross-device Wi-Fi localization
abstract
We investigate the use of linear adaptation on Gaussian processes for Wi-Fi localization in a cross-device setting. We focus on the case where one has a training set collected with one or more devices and a very small labeled adaptation set from the test device. We first present an algorithm to find reliable linear fits between a training set and a small adaptation set by exploiting the Gaussian process assumption that all measurements are spatially correlated. Such regression algorithm is more reliable than total least squares when the adaptation set is very small. Second, we show that the regression misfit can be used to model the additional uncertainty in a linearly adapted map for cross-device Wi-Fi localization. When such uncertainty estimate is used to fuse Gaussian process maps created from one or more training sets and an adaptation set, localization performance can be greatly improved given just a few adaptation samples.
Hsiao-Chieh Yen, Chieh-Chih Wang
IPIN2
2012 M2M gossip: why might we want cars to talk about us?
abstract
What could or should your car be saying about you to other cars or other people on the road? In this paper, we present some preliminary results from a multi-state in-vehicle driver monitoring system and position it with respect to our work in M2M communication. We propose to improve our current motion object tracking algorithm with the addition of a driver state variable, allowing cars to make predictions about other cars' trajectories with information beyond position, velocity and maps. We envision a transportation future where autonomous and semi-autonomous vehicles could be talking about us to our benefit and advanced driver assist would extend beyond the vehicle to a network of connected cars.
Jennifer A. Healey, Chieh-Chih Wang, Andreas Dopfer, Chung-Che Yu
AutomotiveUI2
2012 3D AAM based face alignment under wide angular variations using 2D and 3D data
abstract
Active Appearance Models (AAMs) are widely used to estimate the shape of the face together with its orientation, but AAM approaches tend to fail when the face is under wide angular variations. Although it is feasible to capture the overall 3D face structure using 3D data from range cameras, the locations of facial features are often estimated imprecisely or incorrectly due to depth measurement uncertainty. Face alignment using 2D and 3D images suffer from different issues and have varying reliability in different situations. The existing approaches introduce a weighting function to balance 2D and 3D alignments in which the weighting function is tuned manually and the sensor characteristics are not taken into account. In this paper, we propose to balance 3D face alignment using 2D and 3D data based on the observed data and the sensors characteristics. The feasibility of wide-angle face alignment is demonstrated using two different sets of depth and conventional cameras. The experimental results show that a stable alignment is achieved with a maximum improvement of 26% compared to 3D AAM using 2D image and 30% improvement over the state-of-the-art 3DMM methods in terms of 3D head pose estimation.
Hao-Hsueh Wang, Andreas Dopfer, Chieh-Chih Wang
ICRA3
2011 Vision-based cooperative simultaneous localization and tracking
abstract
Localization is one of the most essential capabilities of autonomous robots. Cooperative localization has been proved to be effective in multi-robot localization. However, nearby moving objects could degrade the cooperative localization performance. In this paper, we demonstrate that the cooperative simultaneous localization and tracking approach is superior in challenging scenarios. Localization and moving object tracking are mutually beneficial. The proposed approach is evaluated using humanoid robots in the RoboCup environment in which only uncertain data from onboard cameras and odometry are used. Ample experimental results with ground truthing from laser scanners demonstrate the accuracy and feasibility of the proposed vision-based cooperative simultaneous localization and tracking algorithm.
Chun-Hua Chang, Shao-Chen Wang, Chieh-Chih Wang
ICRA3
2011 Achieving undelayed initialization in monocular SLAM with generalized objects using velocity estimate-based classification
abstract
Based on the framework of simultaneous localization and mapping (SLAM), SLAM with generalized objects (GO) has an additional structure to allow motion mode learning of generalized objects, and calculates a joint posterior over the robot, stationary objects and moving objects. While the feasibility of monocular SLAM has been demonstrated and undelayed initialization has been achieved using the inverse depth parametrization, it is still challenging to achieve undelayed initialization in monocular SLAM with GO because of the delay decision of static and moving object classification. In this paper, we propose a simple yet effective static and moving object classification method using the velocity estimates directly from SLAM with GO. Compared to the existing approach in which the observations of a new/unclassified feature can not be used in state estimation, the proposed approach makes the uses of all observations without any delay to estimate the whole state vector of SLAM with GO. Both Monte Carlo simulations and real experimental results demonstrate the accuracy of the proposed classification algorithm and the estimates of monocular SLAM with GO.
Chen-Han Hsiao, Chieh-Chih Wang
ICRA2
2011 Feasibility grids for localization and mapping in crowded urban scenes
abstract
Localization and mapping are fundamental tasks in mobile robotics. State-of-the-arts often rely on the static world assumption using the occupancy grids. However, the real environment is typically dynamic. We propose the feasibility grids to facilitate the representation of both the static scene and the moving objects. The dual sensor models are introduced to discriminate between stationary and moving objects in mobile robot localization. Instead of estimating the occupancy states, the feasibility grids maintain the stochastic estimates of the feasibility (crossability) states of the environment. Given that an observation can be decomposed into stationary objects and moving objects, incorporating the feasibility grids in localization yields performance improvements over the occupancy grids, particularly in highly dynamic environments. Our approach is extensively evaluated using real data acquired with a planar laser range finder. The experimental results show that the feasibility grid is capable of rapid convergence and robust performance in mobile robot localization by taking into account moving object information. A root mean squares accuracy of within 50 cm is achieved, without the aid of GPS, which is sufficient for autonomous navigation in crowded urban scenes. The empirical results suggest that the performance of localization can be improved when handling the changing environment explicitly. I.
Shao-Wen Yang, Chieh-Chih Wang
ICRA2
2010 RANSAC matching: Simultaneous registration and segmentation
abstract
The iterative closest points (ICP) algorithm is widely used for ego-motion estimation in robotics, but subject to bias in the presence of outliers. We propose a random sample consensus (RANSAC) based algorithm to simultaneously achieving robust and realtime ego-motion estimation, and multi-scale segmentation in environments with rapid changes. Instead of directly sampling on measurements, RANSAC matching investigates initial estimates at the object level of abstraction for systematic sampling and computational efficiency. A soft segmentation method using a multi-scale representation is exploited to eliminate segmentation errors. By explicitly taking into account the various noise sources degrading the effectiveness of geometric alignment: sensor noise, dynamic objects and data association uncertainty, the uncertainty of a relative pose estimate is calculated under a theoretical investigation of scoring in the RANSAC paradigm. The improved segmentation can also be used as the basis for higher level scene understanding. The effectiveness of our approach is demonstrated qualitatively and quantitatively through extensive experiments using real data.
Shao-Wen Yang, Chieh-Chih Wang, Chun-Hua Chang
ICRA2
2010 Stereo-based simultaneous localization, mapping and moving object tracking
abstract
Vision based simultaneous localization and mapping (SLAM) has recently received much research interest. However, vision based SLAM could be corrupted with the inclusion of moving entities, which makes it hard to operate in dynamic environments. Simultaneous localization, mapping and moving object tracking (SLAMMOT) serves as a solution to deal with moving objects while performing SLAM. The existing work has shown the feasibility of monocular SLAMMOT in dynamic environments. However, monocular SLAMMOT inherits the observability issue of bearings-only tracking in which moving entities would be unobservable according to motions of the camera and moving objects. In this paper, stereo-based SLAMMOT is proposed to solve the observability issue as well as increase the accuracy of localization, mapping and tracking. Simulation and experimental results demonstrate that the proposed stereo SLAMMOT is superior than monocular SLAMMOT in dynamic environments.
Kuen-Han Lin, Chieh-Chih Wang
IROS2
2009 Simultaneous localization of mobile robot and multiple sound sources using microphone array
abstract
Sound source localization is an important function in robot audition. The existing works perform sound source localization using static microphone arrays. This work proposes a framework that simultaneously localizes the mobile robot and multiple sound sources using a microphone array on the robot. First, an eigenstructure-based generalized cross correlation method for estimating time delays between microphones under multi-source environments is described. A method to compute the far field source directions as well as the speed of sound using the estimated time delays is proposed. In addition, the correctness of the sound speed estimate is utilized to eliminate spurious sources, which greatly enhances the robustness of sound source detection. The arrival angles of the detected sound sources are used as observations in a bearings-only SLAM procedure. As the source signals are not persistent and there is no identification of the signal content, data association is unknown which is solved using FastSLAM. The experimental results demonstrate the effectiveness of the proposed approaches.
Jwu-Sheng Hu, Chen-Yu Chan, Cheng-Kang Wang, Chieh-Chih Wang
ICRA4
2009 Multiple-model RANSAC for ego-motion estimation in highly dynamic environments
abstract
Robust ego-motion estimation in urban environments is a key prerequisite for making a robot truly autonomous, but is not easily achievable as there are two motions involved: the motions of moving objects and the motion of the robot itself. We proposed a random sample consensus (RANSAC) based ego-motion estimator to deal with highly dynamic environments using one planar laser scanner. Instead of directly sampling on individual measurements, the RANSAC process is performed at a higher level abstraction for systematic sampling and computational efficiency. We proposed a multiple-model approach to solve the problems of ego-motion estimation and moving object detection jointly in a RANSAC paradigm. To accommodate RANSAC to multiple models - a static environment model for ego-motion estimation and a moving object model for moving object detection, a compact representation models moving object information implicitly is proposed. Moving objects are successfully detected without incorporating any grid maps, that are inherently time and space consuming. The experimental results show that accurate identification of static environments can help classification of moving objects, whereas discrimination of moving objects also yields better ego-motion estimation, particularly in environments containing a significant percentage of moving objects.
Shao-Wen Yang, Chieh-Chih Wang
ICRA2
2008 Dealing with laser scanner failure: Mirrors and windows
abstract
This paper addresses the problem of laser scanner failure on mirrors and windows. Mirrors and glasses are quite common objects that appear in our daily lives. However, while laser scanners play an important role nowadays in the field of robotics, there are very few literatures that address the related issues such as mirror reflection and glass transparency. We introduce a sensor fusion technique to detect the potential obstacles not seen by laser scanners. A laser-based mirror tracker is also proposed to figure out the mirror locations in the environment. The mirror tracking method is seamlessly integrated with the occupancy grid map representation and the mobile robot localization framework. The proposed approaches have been demonstrated using data from sonar sensors and a laser scanner equipped on the NTU-PAL5 robot. Mirrors and windows, as potential obstacles, are successfully detected and tracked.
Shao-Wen Yang, Chieh-Chih Wang
ICRA2
2008 3D active appearance model for aligning faces in 2D images
abstract
Perceiving human faces is one of the most important functions for human robot interaction. The active appearance model (AAM) is a statistical approach that models the shape and texture of a target object. According to a number of the existing works, AAM has a great success in modeling human faces. Unfortunately, the traditional AAM framework could fail when the face pose changes as only 2D information is used to model a 3D object. To overcome this limitation, we propose a 3D AAM framework in which a 3D shape model and an appearance model are used to model human faces. Instead of choosing a proper weighting constant to balance the contributions from appearance similarity and the constraint on consistent 2D shape with 3D shape in the existing work, our approach directly matches 2D visual faces with the 3D shape model. No balancing weighting between 2D shape and 3D shape is needed. In addition, only frontal faces are needed for training and non-frontal faces can be aligned successfully. The experimental results with 20 subjects demonstrate the effectiveness of the proposed approach.
Chun-Wei Chen, Chieh-Chih Wang
IROS2
2007 Interacting Object Tracking in Crowded Urban Areas
abstract
Tracking in crowded urban areas is a daunting task. High crowdedness causes challenging data association problems. Different motion patterns from a wide variety of moving objects make motion modeling difficult. Accompanying with traditional motion modeling techniques, this paper introduces a scene interaction model and a neighboring object interaction model to respectively take long-term and short-term interactions between the tracked objects and its surroundings into account. With the use of the interaction models, anomalous activity recognition is accomplished easily. In addition, move-stop hypothesis tracking is applied to deal with move-stop-move maneuvers. All these approaches are seamlessly inter-graded under the variable-structure multiple-model estimation framework. The proposed approaches have been demonstrated using data from a laser scanner mounted on the PALI robot at a crowded intersection. Interacting pedestrians, bicycles, motorcycles, cars and trucks are successfully tracked in difficult situations with occlusion.
Chieh-Chih Wang, Tzu-Chien Lo, Shao-Wen Yang
ICRA1
2004 A hierarchical object based representation for simultaneous localization and mapping
abstract
Accomplishing simultaneous localization and mapping (SLAM) in very large city environments is a great challenge because of theoretical and practical issues on computational complexity, dynamic environment, representation and data association. In this paper, we describe practical algorithms for dealing with the representation issues. Feature-based, grid-based and direct methods are integrated into the framework of the hierarchical object based representation. The sampling and correlation based range image matching algorithm is developed to tackle the problem arising from uncertain, sparse and featureless data in outdoor environments. Experimental results of a 800 meter /spl times/ 600 meter neighborhood demonstrate the feasibility of city-sized SLAM.
Chieh-Chih Wang, Charles E. Thorpe
IROS1
2003 Online simultaneous localization and mapping with detection and tracking of moving objects: theory and results from a ground vehicle in crowded urban areas
abstract
The simultaneous localization and mapping (SLAM) with detection and tracking of moving objects (DATMO) problem is not only to solve the SLAM problem in dynamic environments but also to detect and track these dynamic objects. In this paper, we derive the Bayesian formula of the SLAM with DATMO problem, which provides a solid basis for understanding and solving this problem. In addition, we provide a practical algorithm for performing DATMO from a moving platform equipped with range sensors. The probabilistic approach to solve the whole problem has been implemented with the Navlab11 vehicle. More than 100 miles of experiments in crowded urban areas indicated that SLAM with DATMO is indeed feasible.
Chieh-Chih Wang, Charles E. Thorpe, Sebastian Thrun
ICRA1
2002 Simultaneous Localization and Mapping with Detection and Tracking of Moving Objects
abstract
Both simultaneous localization and mapping (SLAM) and detection and tracking of moving objects (DTMO) play key roles in robotics and automation. For certain constrained environments, SLAM and DTMO are becoming solved problems, but for robots working outdoors and at high speeds, SLAM and DTMO are still incomplete. In earlier works, SLAM and DTMO are treated as two separate problems. In fact, they can be complementary to one another. In this paper, we present a new method to integrate SLAM and DTMO to solve both problems simultaneously for both indoor and outdoor applications. The results of experiments carried out with CMU Navlab8 and Navlab11 vehicles with the maximum speed of 45 mph in crowded urban and suburban areas verify the described work.
Chieh-Chih Wang, Charles E. Thorpe
ICRA1