Jwu-Sheng Hu

dblp:95/6552 · DBLP profile ↗
← Back
39ranked-venue papers
25as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 21 first-authorSystems, architecture and hardware · 22 · 16 first-authorHuman-computer interaction and ubiquitous computing · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Robot navigation and mapping · 74% 3D vision · 10% Motion planning and robot control · 7%
Computer graphics and multimedia
3 papers
Audio and music processing · 100%
Human-computer interaction and pervasive computing
2 papers
Human-robot interaction · 56% Wearable and physiological sensing · 44%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Computer networks
1 paper
Wireless sensing and localization · 100%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
visual odometry
0.222014
IMU-assisted monocular visual odometry including the human walking model for wearable applications · ICRA 2013
A sliding-window visual-IMU odometer based on tri-focal tensor geometry · ICRA 2014
Robotics › Robot navigation and mapping › localization › visual-inertial navigation
multi-state constraint kalman filter
0.212014
A sliding-window visual-IMU odometer based on tri-focal tensor geometry · ICRA 2014
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry
0.212014
A sliding-window visual-IMU odometer based on tri-focal tensor geometry · ICRA 2014
Robotics › Robot navigation and mapping
localization
0.212013
IMU-assisted monocular visual odometry including the human walking model for wearable applications · ICRA 2013
Robotics › Robot navigation and mapping › visual odometry
monocular visual odometry
0.212013
IMU-assisted monocular visual odometry including the human walking model for wearable applications · ICRA 2013
Robotics › Robot navigation and mapping › sensor fusion
visual-inertial fusion
0.212013
IMU-assisted monocular visual odometry including the human walking model for wearable applications · ICRA 2013
Audio and music processing
beamforming
0.212013
Robust Adaptive Beamformer for Speech Enhancement Using the Second-Order Extended H∞ Filter · IEEE Trans. Speech Audio Process. 2013
Audio and music processing
speech enhancement
0.212013
Robust Adaptive Beamformer for Speech Enhancement Using the Second-Order Extended H∞ Filter · IEEE Trans. Speech Audio Process. 2013
Computer vision › 3D vision
camera calibration
0.112011
Calibration of an eye-to-hand system using a laser pointer on hand and planar constraints · ICRA 2011
Robotics › Motion planning and robot control › robot calibration
hand-eye calibration
0.112011
Calibration of an eye-to-hand system using a laser pointer on hand and planar constraints · ICRA 2011
Audio and music processing › sound source localization
direction-of-arrival estimation
0.112011
Wake-up-word detection for robots using spatial eigenspace consistency and resonant curve similarity · ICRA 2011
Audio and music processing
speech processing
0.112011
Wake-up-word detection for robots using spatial eigenspace consistency and resonant curve similarity · ICRA 2011
Robotics › Robot navigation and mapping › SLAM
bearing-only SLAM
0.112009
Simultaneous localization of mobile robot and multiple sound sources using microphone array · ICRA 2009
Robotics › Robot navigation and mapping
SLAM
0.112009
Simultaneous localization of mobile robot and multiple sound sources using microphone array · ICRA 2009
Robotics › Robot navigation and mapping
sound source localization
0.112009
Simultaneous localization of mobile robot and multiple sound sources using microphone array · ICRA 2009
Wireless sensing and localization
acoustic source localization
0.112009
Location Classification of Nonstationary Sound Sources Using Binaural Room Distribution Patterns · IEEE Trans. Speech Audio Process. 2009
Computer vision › Video understanding and tracking › object tracking › probabilistic tracking
particle filter tracking
0.112007
3-D Human Posture Recognition System Using 2-D Shape Features · ICRA 2007
Computer vision › Face, body and person analysis › human pose estimation
pose detection
0.112007
3-D Human Posture Recognition System Using 2-D Shape Features · ICRA 2007
Computer vision › 3D vision › multi-view geometry › multifocal tensor
trifocal tensor
0.112014
A sliding-window visual-IMU odometer based on tri-focal tensor geometry · ICRA 2014
Mathematical optimization › optimization under uncertainty
robust optimization
0.012013
Robust Adaptive Beamformer for Speech Enhancement Using the Second-Order Extended H∞ Filter · IEEE Trans. Speech Audio Process. 2013
Mathematical optimization › optimization under uncertainty › robust optimization
worst-case optimization
0.012013
Robust Adaptive Beamformer for Speech Enhancement Using the Second-Order Extended H∞ Filter · IEEE Trans. Speech Audio Process. 2013
Computer vision › Image recognition and object detection
shape features
0.012007
3-D Human Posture Recognition System Using 2-D Shape Features · ICRA 2007
Natural language and speech › Speech recognition and synthesis › speech enhancement
microphone array speech enhancement
0.012006
Speaker Attention System for Mobile Robots using Microphone Array and Face Tracking · ICRA 2006

Methods — techniques the papers use, named apart from their topics

unscented kalman filter · 0.3second-order extended h-infinity filter · 0.3kinematic walking model · 0.3kalman filter · 0.3MVDR beamformer · 0.3moving pole model · 0.2gaussian mixture model · 0.2sliding window · 0.2multi-state constraint kalman filter · 0.2RANSAC · 0.2spatial eigenspace consistency · 0.1resonant curve similarity · 0.1nonlinear optimization · 0.1laser pointer · 0.1closed-form solution · 0.1bayes risk detector · 0.1generalized cross correlation · 0.1FastSLAM · 0.1
YearPublicationVenuePosition
2018 A Texture Generation Approach for Detection of Novel Surface Defects
abstract
Surface defect detection is challenging due to varying defect types and their novelties. Because of this, it is hard for algorithms to implement across datasets. Moreover, current automated optical inspection (AOI) machines cannot handle this novelty effectively. In this work, we develop a new method for surface defect detection based on generative models, which can detect novelty according to learned distributions. Experimental results on real industrial datasets show that the proposed method can successfully construct the surface texture pattern generator. By transforming the image through the generator to the corresponding latent space, the defects can be separated effectively without a tedious effort of annotation in a large amount of training data.
Yu-Ting Kevin Lai, Jwu-Sheng Hu
SMC2
2014 A sliding-window visual-IMU odometer based on tri-focal tensor geometry
abstract
This paper presents an odometer architecture which combines a monocular camera and an inertial measurement unit (IMU). The trifocal tensor geometry relationship between three images is used as camera measurement information, which makes the proposed method without estimating the 3D position of feature point. In other words, the proposed method does not have to reconstruct environment. Meanwhile, the camera pose corresponding to each of the three images are refined in filter to form a multi-state constraint Kalman filter (MSCKF). Consequently, this paper proposes a sliding window odometry which has a balance between computational cost and accuracy. Compared with traditional visual odometry or simultaneous localization and mapping (SLAM) method, the proposed method not only meets the requirement of odometer in the ego-motion estimation, but also suit for real-time application. This paper further proposes a random sample consensus (RANSAC) algorithm which is based on three views geometry. The RANSAC algorithm can effectively reject feature points which are mismatch or located on independently moving objects, thus it make the overall algorithm capable of operating in dynamic environment. Experiments are conducted to show the effectiveness of the proposed method in real environment.
Jwu-Sheng Hu, Ming-yuan Chen
ICRA1
2014 Ensuring safety in human-robot coexistence environment
abstract
This paper proposes a safety index and an associated formulation in the optimization-based path planning framework to assess and ensure the safety of human workers in a human-robot coexistence environment. The safety index is evaluated using the ellipsoid coordinates (EC) attached to the robot links that represents the distance between the robot arm and the worker. To account for the inertial effect, the momentum of the robot links are projected onto the coordinates to generate additional measures of safety. The safety index is used as a constraint in the optimization problem so that a collision-free trajectory within a finite time horizon is generated online iteratively for the robot to move towards the desired position. To reduce the computational load for real-time implementation, the formulated optimization problem is further approximated by a quadratic problem. The safety index and the proposed formulations are simulated and validated in a two-link planar robot and the ITRI 7-DoF robot with a human worker moving inside the workspace of the robots.
Chi-Shen Tsai, Jwu-Sheng Hu, Masayoshi Tomizuka
IROS2
2014 Multi-channel post-filtering based on spatial coherence measure
Jwu-Sheng Hu, Ming-Tang Lee
Signal Process.1
2013 IMU-assisted monocular visual odometry including the human walking model for wearable applications
abstract
In this paper, we present a novel approach for monocular visual odometry combined with an inertial measurement unit to estimate the human walking trajectory. The scale factor is computed from the walking speed estimation derived by the kinematic model of human walking. This model not only keeps the important characteristic in biped rolling-foot model, but also makes the speed estimation feasible using human body acceleration. A wearable and calibrated IMU-Camera device is made to mount on the center of human's waist for experimental verification. With the human walking speed and gyroscope measurement, an Unscented Kalman filter is implemented to estimate the current scale factor and refine the attitude of the camera. Experiments in an indoor environment are conducted to show the effectiveness of the proposed method.
Jwu-Sheng Hu, Chin-Yuan Tseng, Ming-yuan Chen, Kuan-Chun Sun
ICRA1
2013 Using UWB sensor for delta robot vibration detection
abstract
This study proposed the ultra-wideband(UWB) sensor to detect the vibration of the delta robot and the distance between UWB sensor and robot arm. Based on the radar propagating principle and the proposed algorithm, the vibration status of robot arm can be obtained. Besides, the vibration status can be reduced by feedback the information to the controller of the robot arm. The advantage of the proposed sensor is that without contacting to the robot arm, UWB sensor is more flexible to be used and will not cause any unnecessary payload as accelerometers. With proper calibration, the simulation of vibration frequency and the real-time absolute position of the robot arm were calculated and demonstrated. The experimental results show that the correlations for measured frequency and distance approximate to 1.
Jyun-Long Chen, Tien-Cheng Tseng, Yu-Yi Cheng, Kuang-I Chang, Jwu-Sheng Hu
IROS5
2013 Robust Adaptive Beamformer for Speech Enhancement Using the Second-Order Extended H∞ Filter
abstract
This paper presents a novel approach to implement the robust minimum variance distortionless response (MVDR) beamformer. The robust MVDR beamformer is based on the optimization of worst-case performance and provides an excellent robustness against an arbitrary but norm-bounded desired signal steering vector mismatch. For real-time consideration, the beamformer was formulated into state-space observer form and the second-order extended (SOE) Kalman filter was derived. However, the SOE Kalman filter assumes an accurate system dynamic and statistics of the noise signals. These assumptions limit the performance under uncertainties. This paper develops the SOEH∞filter for the implementation of the robust MVDR beamformer. The estimation criterion in the SOEH∞filter design is to minimize the worst possible effects of the disturbance signals on the signal estimation errors without a prior knowledge of the disturbance signals statistics. Experimental results demonstrate the performance of the proposed algorithm in a noisy and reverberant environment and show its superiority of the robustness against mismatches over the robust MVDR beamformer based on the SOE Kalman filter.
Jwu-Sheng Hu, Ming-Tang Lee, Chia-Hsing Yang
IEEE Trans. Speech Audio Process.1
2012 Kinematic calibration of manipulator using single laser pointer
abstract
This paper proposes a robot kinematic calibration system including a laser pointer installed on the manipulator, a stationary camera, and a planar surface. The laser pointer beams to the surface, and the camera observes the projected laser spot. The position of the laser spot is computed according to the geometrical relationships of line-plane intersection. The laser spot position is sensitive to slight difference of the end-effector pose due to the extensibility of laser beam. Inaccurate kinematic parameters cause inaccurate calculation of the end-effector pose, and then the laser spot position by the forward estimation is deviated from the one by camera observation. For calibrating the robot kinematics, the optimal solution of kinematic parameters is obtained by minimizing the laser spot position difference between the forward estimation and camera measurement via the nonlinear optimization method. The proposed kinematic calibration system is cost-efficient and flexible for any manipulator. The proposed method is validated by simulation and experiment.
Jwu-Sheng Hu, Jyun-Ji Wang, Yung-Jung Chang
IROS1
2011 Calibration of an eye-to-hand system using a laser pointer on hand and planar constraints
abstract
This work proposes a technique for calibration of an eye-to-hand system. The target of the hand-eye calibration is to estimate the geometric transformation between the hand and the eye. This calibration method further considers camera intrinsic parameters and geometric relations of a working plane in space at the same time. A laser pointer casually mounted on the hand is utilized. By manipulating the robot and projecting the laser beam on a plane of unknown orientations, a batch of related image positions of light-spots are extracted from images of the camera. Since the laser is rigidly mounted and the plane is fixed at each orientation, the geometric parameters and measurement data must obey a certain nonlinear constraints and the solutions of parameters can be estimated accordingly. A close-form solution is developed by decoupling the nonlinear equations into linear forms to compute all of the initial values. As a result, the calibration method does not need any manual initial guess of the unknown parameters. To achieve a higher accuracy, a nonlinear optimization method is implemented to refine the estimation. The advantage of using laser pointer is that this technique can be used for the case when the eye does not see the hand. Experimental results of simulations and real data are presented to show the validity and the simple requirements of the proposed algorithm.
Jwu-Sheng Hu, Yung-Jung Chang
ICRA1
2011 Wake-up-word detection for robots using spatial eigenspace consistency and resonant curve similarity
abstract
In this paper, we propose a method to detect the wake-up-word (WUW) using microphone array for human-robot interaction. The consistency of the spatial eigenspaces formed by the speech source at different frequencies and the resonant curve similarity of the WUW are used as the features for detection. These features are processed and detected separately and the result is determined by cascading individual outcome using Bayes risk detector. This proposed method can keep a high recognition rate under very low signal-to-noise ratio (SNR) conditions. In addition, this method can estimate the direction of arrivals of the sound source, and the proposed architecture is easy to expand by adding detectors with other features in the cascaded manner to further improve the recognition rate.
Jwu-Sheng Hu, Ming-Tang Lee, Ting-Chao Wang
ICRA1
2011 Resource Management for Robotic Applications
abstract
This paper presents Robotic Application Resource Management Services (RARMS), a collection of tools for integrating reusable software components of a wide class of robotic applications on Microsoft Windows, a general-purpose, commodity operating system. RARMS includes (1) a resource allocation tool called RAAPT-HV for partitioning the available processors into a specified number of virtual processors, allocating available resources to virtual processors and assigning robotic software components to virtual processors and (2) a robot-class scheduling service (RC SS) which helps independently developed components prioritize relative to each other in a way that is consistent with their timing requirements. These tools aim to provide time-sensitive components with satisfactory responsiveness in an open environment without serious impact on the performance of other components. To demonstrate the effectiveness of RARMS, we adopted for experimentation and evaluation purposes several commonly-used components of delivery robots, including face detection, speech recognition, video streaming and path planning. The results of our experiments showed that these tools can help to achieve satisfactory performance for these software components of robotic applications.
Yi-Zong Ou, Edward T.-H. Chu, Wen-wei Lu, Jane W.-S. Liu, Ta-Chih Hung, Jwu-Sheng Hu
TrustCom6
2010 A robotic ball catcher with embedded visual servo processor
abstract
In this work we present a robotic ball catcher with embedded visual servo processor. The embedded visual servo processor with powerful parallel computing capability is used as the computation platform to track and triangulate a flying ball's position in 3D based on stereo vision. A recursive least squares algorithm for model-based path prediction of the flying ball is used to determine the catch time and position. Experimental results for real time catching of a flying ball are presented by a 6-DOF robot arm. The percentage of success rate of the robotic ball catcher was found to be approximately 60% for the ball thrown to it from five meters away.
Jwu-Sheng Hu, Ming-Chih Chien, Yung-Jung Chang, Yen-Chung Chang, Shyh-Haur Su, Jwu-Jiun Yang, Chen-Yu Kai
IROS1
2010 A ball-throwing robot with visual feedback
abstract
This work presents a robot system for throwing a ball into a basket. A stereo vision system is used to measure the position of the target in 3D space. The ball-throwing transformation between the input command of the robot system and the target position is built by cubic polynomial. Through ball-throwing transformation with visual feedback for target position, the robot throws the ball toward the target which can randomly move to everywhere in visible field. The percentage of successful ball-throwing for target within three meters was found to be approximately 99%.
Jwu-Sheng Hu, Ming-Chih Chien, Yung-Jung Chang, Shyh-Haur Su, Chen-Yu Kai
IROS1
2010 Speech signal enhancement under multiple interferences using transfer function ratio beamformer
abstract
In many practical environments, the desired speech signal is usually contaminated not only by stationary noise but also nonstationay interferences, such as competing speech. This paper proposes a speech enhancement method which can extract desired speech in a multiple interferences and reverberant environment. The proposed method uses transfer function ratio beamformer and multi-channel adaptive filter algorithm. The virtual sound source concept is proposed to simplify the theoretical treatment for multiple competing speeches. In addition, a transfer function ratio estimation method in a more practical scenario is also proposed. The experiments are performed in a real room acoustic environment.
Jwu-Sheng Hu, Chia-Hsing Yang
IROS1
2010 An investigation of time delay estimation in room acoustic environment using magnitude ratio
Jwu-Sheng Hu, Chia-Hsing Yang, Wei-Han Liu
Signal Process.1
2009 Simultaneous localization of mobile robot and multiple sound sources using microphone array
abstract
Sound source localization is an important function in robot audition. The existing works perform sound source localization using static microphone arrays. This work proposes a framework that simultaneously localizes the mobile robot and multiple sound sources using a microphone array on the robot. First, an eigenstructure-based generalized cross correlation method for estimating time delays between microphones under multi-source environments is described. A method to compute the far field source directions as well as the speed of sound using the estimated time delays is proposed. In addition, the correctness of the sound speed estimate is utilized to eliminate spurious sources, which greatly enhances the robustness of sound source detection. The arrival angles of the detected sound sources are used as observations in a bearings-only SLAM procedure. As the source signals are not persistent and there is no identification of the signal content, data association is unknown which is solved using FastSLAM. The experimental results demonstrate the effectiveness of the proposed approaches.
Jwu-Sheng Hu, Chen-Yu Chan, Cheng-Kang Wang, Chieh-Chih Wang
ICRA1
2009 Self-balancing control and manipulation of a glove puppet robot on a two-wheel mobile platform
abstract
This video shows a continuing work of the glove puppet robot presented before. The major improvement from the previous work is to mount the 9-DOF mechanism, which mimicking a glove puppet manipulation, on a two-wheel mobile platform. The platform provides agile movements of the robot but itself is an unstable system. Hence, a self-balancing controller is implemented by considering the motion as well as configuration variation of the upper body (the 9-DOF mechanism). The control law utilizes the principle of computed torque method with online identification of related parameters using various sensors including an accelerometer. The incline angle is obtained by fusing a gyroscope and a tilt sensor. Under the balancing control, the forward motion of the robot is achieved by giving a desired tilt angle profile. To minimize the footprint of electronics, the controller is implemented using an 8-bit single-chip microcontroller. Further, to enhance the interaction capability of the system, a simple gesture coding using dynamic time warping identification method with Markov model is implemented for the data glove to recognize the puppet gesture by human hand.
Jwu-Sheng Hu, Jyun-Ji Wang, Guan-Cyun Sun
IROS1
2009 Estimation of sound source number and directions under a multi-source environment
abstract
Sound source localization is an important feature in robot audition. This work proposes a sound source number and directions estimation method by using the delay information of microphone array. An eigenstructure-based generalized cross correlation method is proposed to estimate time delay between microphones. Upon obtaining the time delay information, the sound source direction and velocity can be estimated by least square method. In multiple sound source case, the time delay combination among microphones is arranged such that the estimated sound speed value falls within an acceptable range. By accumulating the estimation results of sound source direction and using adaptive K-means++ algorithm, the sound source number and directions can be estimated.
Jwu-Sheng Hu, Chia-Hsing Yang, Cheng-Kang Wang
IROS1
2009 EMWF for Flexible Automation and Assistive Devices
abstract
This paper describes an embedded workflow framework (EMWF) that enables flexible personal and home automation and assistive devices and service and social robots (collectively referred to as SISARL) to be built on workflow architecture The process definition language supported by EMWF is called SISARL-XPDL. It consists of a subset of the WfMC standard XML Process Definition Language (XPDL) 2.0, together with elements that implement common mechanisms for robot behavior coordination. SISARL-XPDL definitions of workflows are first translated into standard XPDL and execution directives and then are parsed either directly into binary workflow scripts for execution or into intermediate scripts in C. EMWF provides workflow engines for Linux and Windows CE platforms. The engines are written in C in order to keep their memory footprint and runtime overhead small. Performance data show that the overheads introduced by the engine and workflow data are tolerable for most SISARL devices.
Ting-Shuo Chou, Su-Ying Chang, Yung-Feng Lu, Yu-Chung Wang, M. K. Ouyang, Chi-Sheng Shih 0001, Tei-Wei Kuo, Jwu-Sheng Hu, Jane W.-S. Liu
IEEE Real-Time and Embedded Technology and Applications Symposium8
2009 Location Classification of Nonstationary Sound Sources Using Binaural Room Distribution Patterns
abstract
This paper discusses the relationships between the nonstationarity of sound sources and the distribution patterns of interaural phase differences (IPDs) and interaural level differences (ILDs) based on short-term frequency analysis. The amplitude variation of nonstationary sound sources is modeled by the exponent of polynomials from the concept of moving pole model. According to the model, the sufficient condition for utilizing the distribution patterns of IPDs and ILDs to localize a nonstationary sound source is suggested and the phenomena of multiple peaks in the distribution pattern can be explained. Simulation is performed to interpret the relation between the distribution patterns of IPD and ILD and the nonstationary sound source. Furthermore, a Gaussian-mixture binaural room distribution model (GMBRDM) is proposed to model distribution patterns of IPDs and ILDs for nonstationary sound source location classification. The effectiveness and performance of the proposed GMBRDM are demonstrated by experimental results.
Jwu-Sheng Hu, Wei-Han Liu
IEEE Trans. Speech Audio Process.1
2008 Adaptive signal blocking for generalized sidelobe canceller using matched filter array
abstract
This work proposes an adaptive beamformer based on generalized sidelobe canceller (GSC) structure with novel blocking matrix design. The classical GSC presented by Griffiths and Jim suffers from desired signal cancellation problems due to the complicated acoustic environment. This work utilizes the pseudo-inverse property of matched filters and subarray structure to design a new blocking matrix of the GSC. For practical implementation, matched filter ratios between microphone pairs are estimated, instead of estimating matched filters from the sound source to each microphone. Simulation results are presented to show the effectiveness of the proposed method.
Jwu-Sheng Hu, Chia-Hsing Yang
ICASSP1
2008 A new spatial-color mean-shift object tracking algorithm with scale and orientation estimation
abstract
In this paper, we propose a new mean-shift tracking algorithm based on a novel similarity measure function. The joint spatial-color feature is used as our basic model elements. The target image is modeled with the kernel density estimation and the new similarity measure functions is developed using the expectation of the estimated kernel density. With these new similarity measure functions, two similarity-based mean-shift tracking algorithms are derived. To enhance the robustness, the weighted background information is added into the proposed tracking algorithm. In order to solve the object deformation problem, the principal component analysis is used to update the orientation of the tracking object, and corresponding eigenvalues are used to monitor the scale of the object. The experimental results show that the new similarity-based tracking algorithms can be implemented in real-time and are able to track the moving object with an automatic update of the orientation and scale.
Chung-Wei Juan, Jwu-Sheng Hu
ICRA2
2008 Calibration and on-line data selection of multiple optical flow sensors for mobile robot localization
abstract
This paper proposes a calibration method as well as a computational algorithm to integrate the data of multiple optical flow sensors for 2-dimensional trajectory measurement. Optical flow sensors offer a different kind of odometer as compared with the wheel encoder. Using multiple sensors, it is possible to reduce the effect of measurement uncertainties. Since all sensors are mounted on a rigid body, their measurement data must obey a certain relation. This relation is utilized in this paper and mathematical formulations are developed to realize the computation. It is shown that the calibration procedure can be cast as an optimization problem given measurement data. Further, the rigid-body relation is formulated as a null-space constraint using the calibrated parameters. During operation, unreliable sensor measurements can be removed by accessing the error distance to the null space. Experimental results are presented to support the proposed methods.
Jwu-Sheng Hu, Yung-Jung Chang, Yu-Lun Hsu
IROS1
2008 The glove puppet robot: X-puppet
abstract
The glove puppet is a traditional art in Taiwan. The master puppeteer manipulates the glove puppets and brings each animated puppet character to life. In this work, we attempt to robotize the glove puppet (called X-puppet) along with three types of manipulation interface including a motion editor, a data glove and a motion capture system. The motion editor is a higher level puppet motion composer that creates and combines sequences of control in a timely fashion. The special designed data glove uses a minimum number of sensors to achieve puppet manipulation. It measures the gesture data and maps to the motion of X-puppet in real-time. The motion capture system can extract the puppet’s motion from the video and then control the X-puppet to simulate the action. The whole system represents an effort to give this traditional art a new style of performance.
Jwu-Sheng Hu, Jyun-Ji Wang, Guan-Qun Sun
IROS1
2008 A new design on multi-modal robotic focus attention
abstract
Human detection and tracking is important for user-friendly human-robot interaction. The robot should be able to find the user autonomously and keep its attention to the user in a human-like manner. In this paper, a design and experimental study of robust human detection and tracking is presented through fusion several modalities of sensory information. The multi-modal interaction design utilizes a combination of visual, audio, and laser scanner data for reliable detection and tracking of an interested user. During tracking motion, obstacle avoidance behavior will be activated any time required to ensure safety. Furthermore, user can further assign the robot to interact with other user by speech command. Experimental results show that the robot can robustly tracks person under complex scenarios.
Chia-How Lin, Chia-Hsing Yang, Cheng-Kang Wang, Kai-Tai Song, Jwu-Sheng Hu
RO-MAN5
2008 Flexible 3D Object Recognition Framework Using 2D Views via a Similarity-Based Aspect-Graph Approach
abstract
This work presents a flexible framework for recognizing 3D objects from 2D views. Similarity-based aspect-graph, which contains a set of aspects and prototypes for these aspects, is employed to represent the database of 3D objects. An incremental database construction method that maximizes the similarity of views in the same aspect and minimizes the similarity of prototypes is proposed as the core of the framework to build and update the aspect-graph using 2D views randomly sampled from a viewing sphere. The proposed framework is evaluated on various object recognition problems, including 3D object recognition, human posture recognition and scene recognition. Shape and color features are employed in different applications with the proposed framework and the top three matching rates show the efficiency of the proposed method.
Jwu-Sheng Hu, Tzung-Min Su
Int. J. Pattern Recognit. Artif. Intell.1
2008 A spatial-color mean-shift object tracking algorithm with scale and orientation estimation
Jwu-Sheng Hu, Chung-Wei Juan, Jyun-Ji Wang
Pattern Recognit. Lett.1
2008 Indoor sound field feature matching for robot's location and orientation detection
Jwu-Sheng Hu, Wei-Han Liu, Chieh-Cheng Cheng
Pattern Recognit. Lett.1
2007 Design methodology and hands-on practices for Embedded Operating Systems
Yu-Lun Huang, Jwu-Sheng Hu
ICPADS2
2007 3-D Human Posture Recognition System Using 2-D Shape Features
abstract
This paper presents an integrated framework for recognizing 3D human posture from 2D images. A flexible combinational algorithm motivated by the novel view expressed by Cyr and Kimia (2004) is proposed to generate the aspects of 3D human postures as the posture prototype using features extracted from the collected 2D images sampled at random intervals from the viewing sphere. Frequency and phase information of the posture are calculated from the Fourier descriptors (FDs) of the sampled points on the posture contour as the main and assistant features to extract the characteristic views as the aspects. Moreover, a modified particle filter is applied to improve the robustness of human posture recognition for continuous monitoring. Experimental trials on synthetic and real sequences have shown the effectiveness of the proposed method.
Jwu-Sheng Hu, Tzung-Min Su, Pei-Ching Lin
ICRA1
2007 Ubiquitous e-Helpers: An UPnP-based home automation platform
abstract
This paper describes a composeable service platform that supports opportunistic collaboration among smart appliances deployed sporadically and incrementally in living and working spaces. This ubiquitous e-Helper's platform is UPnP based and is consisted of a three-level device/service abstraction and a protocol translation proxy. As the first experiment, the platform is used to perform indoor luminance feedback control using wireless sensors and light dimmers. It outperforms similar systems in its responsiveness to dynamic device behaviors, and the modularity in its hardware/software implementation.
John Kar-Kin Zao, Yu-Chih Liu, Ming-Hsiao Yang, Sheng-Kun Li, Wei-Yu Chen, Ching-Wei Chen, Kuo-Chin Huang, Jwu-Sheng Hu, Lun-Chia Kuo
SMC8
2006 Speaker Attention System for Mobile Robots using Microphone Array and Face Tracking
abstract
This paper presents a real-time human-robot interface system (HRIS), which processes both speech and vision information to improve the quality of communication between human and an autonomous mobile robot. The HRIS contains a real-time speech attention system and a real-time face tracking system. In the speech attention system, a microphone-array voice acquisition system has been developed to estimate the direction of speaker and purify the speaker's speech signal in a noisy environment. The developed face tracking system aims to track the speaker's face under illumination variation and react to the face motion. The proposed HRIS can provide a robot with the abilities of finding a speaker's direction, tracking the speaker's face, moving its body to the speaker, focusing its attention to the speaker who is talking to it, and purifying the speaker's speech. The experimental results show that the HRIS not only purifies speech signal with a significant performance, but also tracks a face under illumination variation in real-time
Kai-Tai Song, Jwu-Sheng Hu, Chi-Yi Tsai, Chung-Min Chou, Chieh-Cheng Cheng, Wei-Han Liu, Chia-Hsing Yang
ICRA2
2006 Location and Orientation Detection of Mobile Robots Using Sound Field Features under Complex Environments
abstract
In this paper, the feasibility of utilizing sound for robot's pose detection is investigated, and a novel and robust robot location and orientation detection method based on sound field features for noisy environment is proposed. Unlike traditional methods, the proposed method does not explicitly consider the characteristic of direct path from sound source to microphones, nor attempt to suppress the effect of reverberations and noise signals. Instead, it utilizes the sound field features of a robot at different location and orientation in a normal environment. The sound field feature is captured by using a probability distribution estimation method called Gaussian mixture model (GMM). The experimental results show that this method can detect robot's location and orientation under both line-of-sight and non-line-of-sight conditions using only two microphones and is robust to environmental noise. Moreover, it can also solve the microphones' mismatch problem and can be applied to both near-field and far-field conditions. Since this method can provide global location and orientation detection, it is suitable to fuse with other localization methods to provide initial conditions for reduction of the search effort, or provide the compensation for localizing certain locations that cannot be detected using other localization methods
Jwu-Sheng Hu, Wei-Han Liu, Chieh-Cheng Cheng, Chia-Hsing Yang
IROS1
2006 Robust Background Subtraction with Shadow and Highlight Removal for Indoor Surveillance
abstract
This work describes a new 3D cone-shape illumination model (CSIM) and a robust background subtraction scheme involving shadow and highlight removal for indoor-environmental surveillance. Foreground objects can be precisely extracted for various post-processing procedures such as recognition. Gaussian mixture model (GMM) is applied to construct a color-based probabilistic background model (CBM) that contains the short-term color-based background model (STCBM) and the long-term color-based background model (LTCBM). STCBM and LTCBM are then proposed to build the gradient-based version of the probabilistic background model (GBM) and the CSIM. In the CSIM, a new dynamic cone-shape boundary in the RGB color space is proposed to distinguish pixels among shadow, highlight and foreground. Furthermore, CBM can be used to determine the threshold values of CSIM. A novel scheme combining the CBM, GBM and CSIM is proposed to determine the background. The effectiveness of the proposed method is demonstrated via experiments in a complex indoor environment
Jwu-Sheng Hu, Tzung-Min Su, Shr-Chi Jeng
IROS1
2006 A Robust Speech Enhancement System for Vehicular Applications Using H∞ Adaptive Filtering
abstract
This work proposes a novel and robust adaptive speech enhancement system, which contains both time-domain and frequency-domain beamformers using Hinfinfiltering approach in vehicle environments. A corresponding microphone array data acquisition hardware is also designed and implemented. Traditionally, mutually matched microphones are needed, but this requirement is not practical. To conquer this issue, the proposed system adapts the mismatch dynamics to allow unmatched microphones to be used in an array. Furthermore, to achieve a satisfactory speech recognition performance, the speech recognizer is usually required to be retrained for different vehicle environments due to different noise characteristics and channel effects. The channel effect usually causes the modeling error in a channel recovery process because of the long channel response. The proposed system using the Hinfinfiltering approach, which makes no assumptions about noise and disturbance, is robust to the modeling error. Consequently, the proposed frequency-domain beamformer provides a satisfactory performance without the need to retrain the speech.
Chieh-Cheng Cheng, Wei-Han Liu, Chia-Hsing Yang, Jwu-Sheng Hu
SMC4
2006 Shape Memorization and Recognition of 3D Objects Using a Similarity-Based Aspect-Graph Approach
abstract
This paper presents an integrated framework for recognizing 3D objects from 2D images. A flexible combinational algorithm motivated by the novel view expressed by Cyr and Kimia [1] is proposed to generate the aspects of a 3D object as the object prototype using features extracted from the collected 2D images sampled at random intervals from the viewing sphere. Fourier descriptors of the sampled points on the object contour and point-to-point lengths are calculated as the features and similarity metrics are applied to extract the characteristic views as the aspects. Moreover, the object prototype can be integrated from new collected 2D views. Besides, foreground detection with shadow and highlight removal is used to improve the facility of capturing the explicit object efficiently. The effectiveness of the proposed method is demonstrated by experiments with different rigid objects and human postures.
Tzung-Min Su, Chun-Chi Lin, Pei-Ching Lin, Jwu-Sheng Hu
SMC4
2006 Kannon: Ubiquitous Sensor/Actuator Technologies for Elderly Living and Care: A Multidisciplinary Effort in National Chiao Tung University, Taiwan
abstract
A team of researchers including computer scientists, electrical and control engineers, architects, industrial designers, human factor engineers, and cognitive scientists in the National Chiao Tung University (NCTU), Taiwan, along with their overseas collaborators launched the project Kannon, a multi-disciplinary effort to develop Adaptive Assistive Technologies that can be deployed incrementally into existing private/public spaces and collaborate opportunistically to offer monitoring, assisting, communicating and rejuvenating services to healthy elders. The team combined the state-of-art information, communication and robotic know-how with the activity oriented method for product design and the modular functional approach in modern architecture in order to devise a holistic support for successful aging. This paper presents the philosophy, approach and first fruits of this project.
John Kar-Kin Zao, Jwu-Sheng Hu, Jin-Chern Chiou, Yu-Lun Huang, Shu-Chen Li, Zee-Yih Kuo, Ming-Chuen Chuang, Shang Hwa Hsu, Yu-Chee Tseng, Jane W.-S. Liu, Chin-Teng Lin
SMC2
2006 Robust Speaker's Location Detection in a Vehicle Environment Using GMM Models
abstract
Abstract-Human-computer interaction (HCI) using speech communication is becoming increasingly important, especially in driving where safety is the primary concern. Knowing the speaker's location (i.e., speaker localization) not only improves the enhancement results of a corrupted signal, but also provides assistance to speaker identification. Since conventional speech localization algorithms suffer from the uncertainties of environmental complexity and noise, as well as from the microphone mismatch problem, they are frequently not robust in practice. Without a high reliability, the acceptance of speech-based HCI would never be realized. This work presents a novel speaker's location detection method and demonstrates high accuracy within a vehicle cabinet using a single linear microphone array. The proposed approach utilize Gaussian mixture models (GMM) to model the distributions of the phase differences among the microphones caused by the complex characteristic of room acoustic and microphone mismatch. The model can be applied both in near-field and far-field situations in a noisy environment. The individual Gaussian component of a GMM represents some general location-dependent but content and speaker-independent phase difference distributions. Moreover, the scheme performs well not only in nonline-of-sight cases, but also when the speakers are aligned toward the microphone array but at difference distances from it. This strong performance can be achieved by exploiting the fact that the phase difference distributions at different locations are distinguishable in the environment of a car. The experimental results also show that the proposed method outperforms the conventional multiple signal classification method (MUSIC) technique at various SNRs.
Jwu-Sheng Hu, Chieh-Cheng Cheng, Wei-Han Liu
IEEE Trans. Syst. Man Cybern. Part B1
2001 Optimal synthesis of a fractional delay FIR filter in a reproducing kernel Hilbert space
abstract
Based on a bandlimited signal model, the optimal fractional delay finite impulse response (FIR) filter and corresponding interpolating error bound is derived in a reproducing kernel Hilbert space. The resulting optimal filtering is a projection onto a prescribed finite dimensional subspace. The connection of this filtering accuracy to the delay time and the filter order is investigated via error analysis.
Shiang-Hwua Yu, Jwu-Sheng Hu
IEEE Signal Process. Lett.2