Katsushi Ikeuchi

dblp:44/4771 · DBLP profile ↗
← Back
287ranked-venue papers
32as first author
10since 2021 · last 2025
0000-0001-9758-9357ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 237 · 26 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 121 · 11 first-author · 1 since 2021Systems, architecture and hardware · 79 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
abstract
In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human’s activity, urgency of the interaction, etc.). We test whether large language models (LLM) and vision language models (VLM) can provide solutions to this problem. We compare four different system-design patterns using LLMs and VLMs, and test on a test set containing 84 human-robot situations. The test set mixes several publicly available datasets and also includes situations where the appropriate action to take is open-ended. Our results using the GPT-4o and Phi-3 Vision model indicate that LLMs and VLMs are capable of handling interaction beginnings when the desired actions are clear. The design using direct image input scored an 89% accuracy on the test set. Of the designs using indirect input, a combined text about human activity and gaze performed best with a 90% accuracy. However, challenges remain in the open-ended situations where the model must choose the priority between the human and robot situation. The design using direct image input mostly prioritized the robot situation, whereas the design with best performance using indirect input mostly prioritized the human situation. Such one-sided behavior could be crucial for practical HRI applications.
Kazuhiro Sasabuchi, Naoki Wake, Atsushi Kanehira, Jun Takamatsu, Katsushi Ikeuchi
RO-MAN5
2025 Plan-and-Act using Large Language Models for Interactive Agreement
abstract
Recent large language models (LLMs) are capable of planning robot actions. In this paper, we explore how LLMs can be used for planning actions with tasks involving situational human-robot interaction (HRI). A key problem of applying LLMs in situational HRI is balancing between "respecting the current human’s activity" and "prioritizing the robot’s task," as well as understanding the timing of when to use the LLM to generate an action plan. In this paper, we propose a necessary plan-and-act skill design to solve the above problems. We show that a critical factor for enabling a robot to switch between passive / active interaction behavior is to provide the LLM with an action text about the current robot’s action. We also show that a second-stage question to the LLM (about the next timing to call the LLM) is necessary for planning actions at an appropriate timing. The skill design is applied to an Engage skill and is tested on four distinct interaction scenarios. We show that by using the skill design, LLMs can be leveraged to easily scale to different HRI scenarios with a reasonable success rate reaching 90% on the test scenarios.
Kazuhiro Sasabuchi, Naoki Wake, Atsushi Kanehira, Jun Takamatsu, Katsushi Ikeuchi
RO-MAN5
2024 LiDAR-camera Calibration using Intensity Variance Cost
abstract
We propose an extrinsic calibration method for LiDAR-camera fusion systems using variations in intensities projected from camera images to the LiDAR point cloud. As the input, the proposed method uses a sequence of LiDAR data and camera images captured while moving the system. Once the camera motion is calculated, camera images are projected onto the point cloud. The variations in the projected intensities at each point are large in the presence of errors in the estimated motion or calibration parameters. Consequently, the extrinsic parameters are optimized for cost minimization based on the intensity variance. In addition, a suitable geometry is proposed for the calibration and verified using simulations. Our experimental results showed that the proposed method accurately performed calibrations using a camera and a sparse multi-beam LiDAR or one-dimensional LiDAR.
Ryoichi Ishikawa, Yoshihiro Sato, Takeshi Oishi, Katsushi Ikeuchi
ICRA5
2023 Text-driven object affordance for guiding grasp-type recognition in multimodal robot teaching
Naoki Wake, Daichi Saito, Kazuhiro Sasabuchi, Hideki Koike, Katsushi Ikeuchi
Mach. Vis. Appl.5
2022 Deep Gesture Generation for Social Robots Using Type-Specific Libraries
abstract
Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the discussion. In the field of robotics, giving conversational agents (humanoid robots or virtual avatars) the ability to properly use gestures is critical, yet remain a task of extraordinary difficulty. This is because given only a text as input, there are many possibilities and ambiguities to generate an appropriate gesture. Different to previous works we propose a new method that explicitly takes into account the gesture types to reduce these ambiguities and generate human-like conversational gestures. Key to our proposed system is a new gesture database built on the TED dataset that allows us to map a word to one of three types of gestures: “Imagistic” gestures, which express the content of the speech, “Beat” gestures, which emphasize words, and “No gestures.” We propose a system that first maps the words in the input text to their corresponding gesture type, generate type-specific gestures and combine the generated gestures into one final smooth gesture. In our comparative experiments, the effectiveness of the proposed method was confirmed in user studies for both avatar and humanoid robot.
Hitoshi Teshima, Naoki Wake, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi
IROS6
2022 Editorial for Special Issue on Computer Vision in the Wild
Cha Zhang, Katsushi Ikeuchi
Int. J. Comput. Vis.3
2022 Kyushu Decorative Tumuli Project: From e-Heritage to Cyber-Archaeology
abstract
Abstract Digitization of cultural assets has become an important sub-area of computer vision (CV). Thus far, the value of digitization has been emphasized in terms of asset preservation and exhibition. The third aspect of digitization value is that the obtained digital data can be used to perform archaeological analysis based on physics and optics theories and simulations. This position paper emphasizes the importance of this third aspect, using our Kyushu decorative tumuli project as an illustrative example. In particular, we focus on the photometric approaches in the third aspect and explain the equipment and methods developed there as well as archaeological findings. This paper, then, proposes to establish this area as “cyber-archaeology” through categorizing and organizing those methodologies.
Katsushi Ikeuchi, Tetsuro Morimoto, Mawo Kamakura, Nobuaki Kuchitsu, Kazutaka Kawano, Tomoo Ikeda
Int. J. Comput. Vis.1
2021 PoseRN: A 2D Pose Refinement Network For Bias-Free Multi-View 3D Human Pose Estimation
abstract
We propose a new 2D pose refinement network that learns to predict the human bias in the estimated 2D pose. There are biases in 2D pose estimations that are due to differences between annotations of 2D joint locations based on annotators’ perception and those defined by motion capture (MoCap) systems. These biases are crafted into publicly available 2D pose datasets and cannot be removed with existing error reduction approaches. Our proposed pose refinement network allows us to efficiently remove the human bias in the estimated 2D poses and achieve highly accurate multi-view 3D human pose estimation.
Akihiko Sayo, Diego Thomas, Hiroshi Kawasaki, Yuta Nakashima, Katsushi Ikeuchi
ICIP5
2021 Verbal Focus-of-Attention System for Learning-from-Observation
abstract
The learning-from-observation (LfO) framework aims to map human demonstrations to a robot to reduce programming effort. To this end, an LfO system encodes a human demonstration into a series of execution units for a robot, which are referred to as task models. Although previous research has proposed successful task-model encoders, there has been little discussion on how to guide a task-model encoder in a scene with spatio-temporal noises, such as cluttered objects or unrelated human body movements. Inspired by the function of verbal instructions guiding an observer’s visual attention, we propose a verbal focus-of-attention (FoA) system (i.e., spatiotemporal filters) to guide a task-model encoder. For object manipulation, the system first recognizes the name of a target object and its attributes from verbal instructions. The information serves as a where-to-look FoA filter to confine the areas in which the target object existed in the demonstration. The system then detects the timings of grasp and release that occurred in the filtered areas. The timings serve as a when-to-look FoA filter to confine the period of object manipulation. Finally, a task-model encoder recognizes the task models by employing the FoA filters. We demonstrate the robustness of the verbal FoA in attenuating spatio-temporal noises by comparing it with an existing action localization network. The contributions of this study are as follows: (1) to propose a verbal FoA for LfO, (2) to design an algorithm to calculate FoA filters from verbal input, and (3) to demonstrate the effectiveness of a verbal FoA in localizing an action by comparing it with a state-of-the-art vision system.
Naoki Wake, Iori Yanokura, Kazuhiro Sasabuchi, Katsushi Ikeuchi
ICRA4
2021 Reconstruction of Geometric and Optical Parameters of Non-Planar Objects with Thin Film
abstract
Here, we propose a novel method to estimate the parameters of non-planar objects with thin film surfaces. Being able to estimate the optical parameters of objects with thin film surfaces has a wide range of applications from industrial inspections to biological and archaeology research. However, there are many challenging issues that need to be overcome to model such parameters. The appearance of thin film objects is highly dependent on the surface orientation and optical parameters such as the refractive index and film thickness. First, we therefore analyzed the optical parameters of non-planar objects with thin film surfaces. Next, we proposed and implemented an analysis procedure and demonstrated its effectiveness for studying planar objects with thin film surfaces. Finally, we developed a device to acquire the shapes and optical parameters of objects with thin film surfaces using a camera and demonstrated the effectiveness of our method experimentally. Then, we surveyed the errors caused by the light source. We discussed the difference between the theoretically obtained parameters and experimental data obtained using a hyper spectral camera.
Yoshie Kobayashi, Tetsuro Morimoto, Imari Sato, Yasuhiro Mukaigawa, Takao Tomono, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 Human Shape Reconstruction with Loose Clothes from Partially Observed Data by Pose Specific Deformation
Akihiko Sayo, Hayato Onizuka, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi
PSIVT6
2019 Dynamic Calibration between a Mobile Robot and SLAM Device for Navigation
abstract
In this paper, we propose a dynamic calibration between a mobile robot and a device using simultaneous localization and mapping (SLAM) technology, which we termed as the SLAM device, for a robot navigation system. The navigation framework assumes loose mounting of SLAM device for easy use and requires an online adjustment to remove localization errors. The online adjustment method dynamically corrects not only the calibration errors between the SLAM device and the part of the robot to which the device is attached but also the robot encoder errors by calibrating the whole body of the robot. The online adjustment assumes that the information of the external environment and shape information of the robot are consistent. In addition to the online adjustment, we also present an offline calibration between a robot and device. The offline calibration is motion-based and we clarify the most efficient method based on the number of degrees-of-freedom of the robot movement. Our method can be easily used for various types of robots with sufficiently precise localization for navigation. In the experiments, we confirm the parameters obtained via two types of offline calibration based on the degree of freedom of robot movement. We also validate the effectiveness of the online adjustment method by plotting localized position errors during a robots intense movement. Finally, we demonstrate the navigation using a SLAM device.
Ryoichi Ishikawa, Takeshi Oishi, Katsushi Ikeuchi
RO-MAN3
2018 Representing a Partially Observed Non-Rigid 3D Human Using Eigen-Texture and Eigen-Deformation
abstract
Reconstruction of the shape and motion of humans from RGB-D is a challenging problem, receiving much attention in recent years. Recent approaches for full-body reconstruction use a statistic shape model, which is built upon accurate full-body scans of people in skin-tight clothes, to complete invisible parts due to occlusion. Such a statistic model may still be fit to an RGB-D measurement with loose clothes but cannot describe its deformations, such as clothing wrinkles. Observed surfaces may be reconstructed precisely from actual measurements, while we have no cues for unobserved surfaces. For full-body reconstruction with loose clothes, we propose to use lower dimensional embeddings of texture and deformation referred to as eigen-texturing and eigen-deformation, to reproduce views of even unobserved surfaces. Provided a full-body reconstruction from a sequence of partial measurements as 3D meshes, the texture and deformation of each triangle are then embedded using eigen-decomposition. Combined with neural-network-based coefficient regression, our method synthesizes the texture and deformation from arbitrary viewpoints. We evaluate our method using simulated data and visually demonstrate how our method works on real data.
Ryosuke Kimura, Akihiko Sayo, Fabian Lorenzo Dayrit, Yuta Nakashima, Hiroshi Kawasaki, Ambrosio Blanco, Katsushi Ikeuchi
ICPR7
2018 LiDAR and Camera Calibration Using Motions Estimated by Sensor Fusion Odometry
abstract
This paper proposes a targetless and automatic camera-LiDAR calibration method. Our approach extends the hand-eye calibration framework to 2D-3D calibration. The scaled camera motions are accurately calculated using a sensor-fusion odometry method. We also clarify the suitable motions for our calibration method. Whereas other calibrations require the LiDAR reflectance data and an initial extrinsic parameter, the proposed method requires only the three-dimensional point cloud and the camera image. The effectiveness of the method is demonstrated in experiments using several sensor configurations in indoor and outdoor scenes. Our method achieved higher accuracy than comparable state-of-the-art methods.
Ryoichi Ishikawa, Takeshi Oishi, Katsushi Ikeuchi
IROS3
2018 Describing Upper-Body Motions Based on Labanotation for Learning-from-Observation Robots
abstract
We have been developing a paradigm that we call learning-from-observation for a robot to automatically acquire a robot program to conduct a series of operations, or for a robot to understand what to do, through observing humans performing the same operations. Since a simple mimicking method to repeat exact joint angles or exact end-effector trajectories does not work well because of the kinematic and dynamic differences between a human and a robot, the proposed method employs intermediate symbolic representations, tasks, for conceptually representing what-to-do through observation. These tasks are subsequently mapped to appropriate robot operations depending on the robot hardware. In the present work, task models for upper-body operations of humanoid robots are presented, which are designed on the basis of Labanotation. Given a series of human operations, we first analyze the upper-body motions and extract certain fixed poses from key frames. These key poses are translated into tasks represented by Labanotation symbols. Then, a robot performs the operations corresponding to those task models. Because tasks based on Labanotation are independent of robot hardware, different robots can share the same observation module, and only different task-mapping modules specific to robot hardware are required. The system was implemented and demonstrated that three different robots can automatically mimic human upper-body operations with a satisfactory level of resemblance.
Katsushi Ikeuchi, Zhaoyuan Ma, Zengqiang Yan, Shunsuke Kudoh, Minako Nakamura
Int. J. Comput. Vis.1
2018 Multiview Rectification of Folded Documents
abstract
Digitally unwrapping images of paper sheets is crucial for accurate document scanning and text recognition. This paper presents a method for automatically rectifying curved or folded paper sheets from a few images captured from multiple viewpoints. Prior methods either need expensive 3D scanners or model deformable surfaces using over-simplified parametric representations. In contrast, our method uses regular images and is based on general developable surface models that can represent a wide variety of paper deformations. Our main contribution is a new robust rectification method based on ridge-aware 3D reconstruction of a paper sheet and unwrapping the reconstructed surface using properties of developable surfaces via conformal mapping. We present results on several examples including book pages, folded letters and shopping receipts.
Shaodi You, Yasuyuki Matsushita, Sudipta N. Sinha, Yusuke Bou, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Realtime Novel View Synthesis with Eigen-Texture Regression
Yuta Nakashima, Fumio Okura, Norihiko Kawai, Ryosuke Kimura, Hiroshi Kawasaki, Katsushi Ikeuchi, Ambrosio Blanco
BMVC6
2017 Radiometric Calibration from Faces in Images
abstract
We present a method for radiometric calibration of cameras from a single image that contains a human face. This technique takes advantage of a low-rank property that exists among certain skin albedo gradients because of the pigments within the skin. This property becomes distorted in images that are captured with a non-linear camera response function, and we perform radiometric calibration by solving for the inverse response function that best restores this low-rank property in an image. Although this work makes use of the color properties of skin pigments, we show that this calibration is unaffected by the color of scene illumination or the sensitivities of the cameras color filters. Our experiments validate this approach on a variety of images containing human faces, and show that faces can provide an important source of calibration data in images where existing radiometric calibration techniques perform poorly.
Chen Li 0031, Stephen Lin 0001, Kun Zhou 0001, Katsushi Ikeuchi
CVPR4
2017 Specular Highlight Removal in Facial Images
abstract
We present a method for removing specular highlight reflections in facial images that may contain varying illumination colors. This is accurately achieved through the use of physical and statistical properties of human skin and faces. We employ a melanin and hemoglobin based model to represent the diffuse color variations in facial skin, and utilize this model to constrain the highlight removal solution in a manner that is effective even for partially saturated pixels. The removal of highlights is further facilitated through estimation of directionally variant illumination colors over the face, which is done while taking advantage of a statistically-based approximation of facial geometry. An important practical feature of the proposed method is that the skin color model is utilized in a way that does not require color calibration of the camera. Moreover, this approach does not require assumptions commonly needed in previous highlight removal techniques, such as uniform illumination color or piecewise-constant surface colors. We validate this technique through comparisons to existing methods for removing specular highlights.
Chen Li 0031, Stephen Lin 0001, Kun Zhou 0001, Katsushi Ikeuchi
CVPR4
2017 ReMagicMirror: Action Learning Using Human Reenactment with the Mirror Metaphor
Fabian Lorenzo Dayrit, Ryosuke Kimura, Yuta Nakashima, Ambrosio Blanco, Hiroshi Kawasaki, Katsushi Ikeuchi, Tomokazu Sato, Naokazu Yokoya
MMM (1)6
2017 Guest Editorial: Best Papers from ICCV 2015
Katsushi Ikeuchi, Christoph Schnörr, Josef Sivic, René Vidal
Int. J. Comput. Vis.1
2016 A 3D Reconstruction with High Density and Accuracy Using Laser Profiler and Camera Fusion System on a Rover
abstract
3D Sensing systems mounted on mobile platform are emerging and have been developed for various applications. In this paper, we propose a profiler scanning system mounted on a rover to scan and reconstruct a bas-relief with high density and accuracy. Our hardware system consists of an omnidirectional camera and a 3D laser scanner. Our method selects good projection points for tracking to estimate motion stably and reject mismatches caused by difference between the positions of laser scanner and camera using an error metric based on the distance from omnidirectional camera to scanned point. We demonstrate that our results has better accuracy than comparable approach. In addition to local motion estimation method, we propose global poses refinement method using multi modal 2D-3D registration and our result shows good consistency between reflectance image and 2D RGB image.
Ryoichi Ishikawa, Menandro Roxas, Yoshihiro Sato, Takeshi Oishi, Takeshi Masuda 0001, Katsushi Ikeuchi
3DV6
2016 e-Intangible Heritage
Katsushi Ikeuchi
BMVC1
2016 Reconstructing Shapes and Appearances of Thin Film Objects Using RGB Images
abstract
Reconstruction of shapes and appearances of thin film objects can be applied to many fields such as industrial inspection, biological analysis, and archaeologic research. However, it comes with many challenging issues because the appearances of thin film can change dramatically depending on view and light directions. The appearance is deeply dependent on not only the shapes but also the optical parameters of thin film. In this paper, we propose a novel method to estimate shapes and film thickness. First, we narrow down candidates of zenith angle by degree of polarization and determine it by the intensity of thin film which increases monotonically along the zenith angle. Second, we determine azimuth angle from occluding boundaries. Finally, we estimate the film thickness by comparing a look-up table of color along the thickness and zenith angle with captured images. We experimentally evaluated the accuracy of estimated shapes and appearances and found that our proposed method is effective.
Yoshie Kobayashi, Tetsuro Morimoto, Imari Sato, Yasuhiro Mukaigawa, Takao Tomono, Katsushi Ikeuchi
CVPR6
2016 Outdoor omnidirectional video completion via depth estimation by motion analysis
abstract
Video completion aims to track, remove, and fill in unwanted regions (holes) of a video sequence. Holes have to be filled-in consistently to create a visually pleasant video output. Challenges arise when big holes propagate along several frames (large spatiotemporal holes) in outdoor videos with variant illumination and structured background. In those cases even forefront video completion approaches based on optical flow fail to complete the holes correctly as 3D information is required to keep the structure of the scene and a wider field of view is needed to handle the large spatiotemporal holes. To overcome these limitations, we propose a novel omnidirectional video completion framework based on depth estimation. First, we recover the depth of the scene from a pixel motion model constrained by known camera pose. The depth map is further improved by a structure-aware refinement. The refined depth map is then employed for color propagation into the holes. We perform a set of experiments to evaluate our approaches for preliminary depth recovery, depth refinement, and color propagation. Our results confirm that the proposed framework generates accurate preliminary depth maps, improves the depth quality maintaining the structure of the scene, and outperforms state-of-the-art optical-flow-based video completion approach in terms of accuracy and visual appeal.
Carlos Morales, Menandro Roxas, Yasuhide Okamoto, Shintaro Ono, Takeshi Oishi, Katsushi Ikeuchi
ICPR6
2016 Perceptual enhancement for stereoscopic videos based on horopter consistency
abstract
Audience discomfort, such as eye strain and dizziness, is one of the urgent issues that virtual reality and 3D movie technologies should tackle. Except for inappropriate horizontal and vertical disparity, one major problem is that people's binocular vergence and focal length in the cinema remain inconsistent from normal visual habits. Psychologists discovered the horopter and Panum's fusional area to describe zero-disparity points projected on the retinas based on accommodation-convergence consistency. In this paper, inspired by these concepts, we propose a stereoscopic effect correction system for perceptual enhancement according to fixated region and scene information. As a preprocessing step, tracking and stereo matching algorithms are implemented to prepare cues for further transformation in 3D space. Then in order to accomplish certain visual effects, we describe a geometric framework for disparity refinement and image warping based on parameter adjustment of the virtual stereoscopic rig. For evaluation, subjective experiments have been conducted to prove the effectiveness of our method. Therefore, our work provides a possibility to improve the audience experience from a formerly underexplored perspective.
Zeyu Wang 0003, Xiaohan Jin, Renju Li, Hongbin Zha, Katsushi Ikeuchi
VRST6
2016 Special issue on IAPR MVA2013 best papers
Masaki Suwa, Yoichi Sato 0001, Katsushi Ikeuchi
Mach. Vis. Appl.3
2016 Adherent Raindrop Modeling, Detectionand Removal in Video
abstract
Raindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Modeling, detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. To accomplish the idea, we first model adherent raindrops using law of physics, and detect raindrops based on these models in combination with motion and intensity temporal derivatives of the input video. Having detected the raindrops, we remove them and restore the images based on an analysis that some areas of raindrops completely occludes the scene, and some other areas occlude only partially. For partially occluding areas, we restore them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity derivative. For completely occluding areas, we recover them by using a video completion technique. Experimental results using various real videos show the effectiveness of our method.
Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 A New Flying Range Sensor: Aerial Scan in Omni-Directions
abstract
This paper presents a new flying sensor system to capture 3D data aerially. The hardware system, consisting of a omni-directional laser scanner and a panoramic camera, can be mounted under a mobile platform (e.g., a balloon or a crane) to achieve the aerial scanning with high resolution and accuracy. Since the laser scanner often requires several minutes to complete an omni-directional scan, the raw data is distorted seriously due to the unknown and uncontrollable movement during the scanning period. To overcome this problem, 1) we first synchronize the two sensors and spherically calibrate them together, 2) our approach then recovers the sensor motion by utilizing the spacial and temporal features extracted both from the image sequences and point clouds, and 3) finally the distorted scans can be rectified with the estimated motion and aligned together automatically. In experiments, we demonstrate that the method achieves a substantially good performance for indoor/outdoor aerial scanning in the applications such as Angkor Wat 3D preservation and manufacturing 3D survey with respect to other state-of-the-art methods.
Bo Zheng 0001, Xiangqi Huang, Ryoichi Ishikawa, Takeshi Oishi, Katsushi Ikeuchi
3DV5
2015 Motion generation of the humanoid robot for teleoperation by task model
abstract
In recent years, the research of humanoid robots that replace human tasks in emergency situations have been widely studied. Currently, many approaches are automate dedicated hardware for each mission. But, at the environment where situation changes, operation by humanoid robot is effective to operate equipments which designed for human. Ultimately, automation is ideal, but under the present circumstances, teleoperation of humanoid robot is effective for corresponding changes of situation. An intuitive interface is required for effectively controlling the humanoid robot from a distant place. Recently, the interfaces that map the human motion to the humanoid robot have become popular because of the development of the motion recognition systems. However, the humanoid robot and human beings have different joint structure, physical ability and weight balance. It is not practical to map the motion directly. There is also the issue of time delay between the operator and the robot. Therefore, it is desirable that the operator performs global judgments and the robot runs semi-autonomously in the local environment. In this paper we propose a method to remotely operate the humanoid robot by the task model. Our method describes human behavior abstractly by the task model and mapped this abstract expressions to humanoid robots, and overcome difference of structure of body. In this work, we operate lever of buggy-type vehicles as a example of mapping using the task model.
Masaya Ogawa, Katsuya Honda, Yoshihiro Sato, Shunsuke Kudoh, Takeshi Oishi, Katsushi Ikeuchi
RO-MAN6
2015 Scene Understanding by Reasoning Stability and Safety
Bo Zheng 0001, Yibiao Zhao, Joey C. Yu, Katsushi Ikeuchi, Song-Chun Zhu
Int. J. Comput. Vis.4
2014 Photometric Stereo Using Internet Images
abstract
Photometric stereo using unorganized Internet images is very challenging, because the input images are captured under unknown general illuminations, with uncontrolled cameras. We propose to solve this difficult problem by a simple yet effective approach that makes use of a coarse shape prior. The shape prior is obtained from multi-view stereo and will be useful in twofold: resolving the shape-light ambiguity in uncalibrated photometric stereo and guiding the estimated normals to produce the high quality 3D surface. By assuming the surface albedo is not highly contrasted, we also propose a novel linear approximation of the nonlinear camera responses with our normal estimation algorithm. We evaluate our method using synthetic data and demonstrate the surface improvement on real data over multi-view stereo results.
Boxin Shi, Kenji Inose, Yasuyuki Matsushita, Ping Tan 0002, Sai-Kit Yeung, Katsushi Ikeuchi
3DV6
2014 Reconstructing Shape and Appearance of Thin Film Objects with Hyper Spectral Sensor
Yoshie Kobayashi, Tetsuro Morimoto, Imari Sato, Yasuhiro Mukaigawa, Katsushi Ikeuchi
ACCV (4)5
2014 Raindrop Detection and Removal from Long Range Trajectories
Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi
ACCV (2)5
2014 Robust 3D Features for Matching between Distorted Range Scans Captured by Moving Systems
abstract
Laser range sensors are often demanded to mount on a moving platform for achieving the good efficiency of 3D reconstruction. However, such moving systems often suffer from the difficulty of matching the distorted range scans. In this paper, we propose novel 3D features which can be robustly extracted and matched even for the distorted 3D surface captured by a moving system. Our feature extraction employs Morse theory to construct Morse functions which capture the critical points approximately invariant to the 3D surface distortion. Then for each critical point, we extract support regions with the maximally stable region defined by extremal region or disconnectivity. Our feature description is designed as two steps: 1) we normalize the detected local regions to canonical shapes for robust matching, 2) we encode each key point with multiple vectors at different Morse function values. In experiments, we demonstrate that the proposed 3D features achieve substantially better performance for distorted surface matching than the state-of-the-art methods.
Xiangqi Huang, Bo Zheng 0001, Takeshi Masuda 0001, Katsushi Ikeuchi
CVPR4
2014 Simultaneous deblur and super-resolution technique for video sequence captured by hand-held video camera
abstract
Nowadays, video camera is commonly used everywhere and demand of retrieving a single shot from video sequence is increasing. Since resolution of video camera is usually lower than that of digital camera, simply cutting out a frame from a video sequence ends up with low quality. Further, because of the necessity of high fps on video camera, video data inevitably contains motion blur and it leads mis-registration between frames which is critical for multi-frame superresolution. In this paper, we propose a method to restore high-resolution image from a video sequence considering motion blur. Since the frame-rate of a video camera is high, motion of the object in successive frames is small, and thus, stable feature tracking during short sequences is possible even if there is a blur. Thus, we adopt a division/integration approach to realize robust tracking for long sequence. We also propose a simultaneous deblur and super-resolution technique using multiple images based on MAP estimation. Experimental results are shown to prove the strength of our method.
Yuki Matsushita, Hiroshi Kawasaki, Shintaro Ono, Katsushi Ikeuchi
ICIP4
2014 Detecting potential falling objects by inferring human action and natural disturbance
abstract
Detecting potential dangers in the environment is a fundamental ability of living beings. In order to endure such ability to a robot, this paper presents an algorithm for detecting potential falling objects, i.e. physically unsafe objects, given an input of 3D point clouds captured by the range sensors. We formulate the falling risk as a probability or a potential that an object may fall given human action or certain natural disturbances, such as earthquake and wind. Our approach differs from traditional object detection paradigm, it first infers hidden and situated “causes (disturbance) of the scene, and then introduces intuitive physical mechanics to predict possible “effects (falls) as consequences of the causes. In particular, we infer a disturbance field by making use of motion capture data as a rich source of common human pose movement. We show that, by applying various disturbance fields, our model achieves a human level recognition rate of potential falling objects on a dataset of challenging and realistic indoor scenes.
Bo Zheng 0001, Yibiao Zhao, Joey C. Yu, Katsushi Ikeuchi, Song-Chun Zhu
ICRA4
2014 Extraction of person-specific motion style based on a task model and imitation by humanoid robot
abstract
In this paper, we present a humanoid robot which extracts and imitates the person-specific differences in motions, which we will call style. Synthesizing human-like and stylistic motion variations according to specific scenarios is becoming important for entertainment robots, and imitation of styles is one variation which makes robots more amiable. Our approach extends a learning from observation (LFO) paradigm which enables robots to understand what a human is doing and to extract reusable essences to be learned. The focus is on styles in the domain of LFO and the representation of them using the reusable essences. In this paper, we design an abstract model of a target motion defined in LFO, observe human demonstrations through the model, and formulate the representation of styles in the context of LFO. Then we introduce a framework of generating robot motions that reflect styles which are automatically extracted from human demonstrations. To verify our proposed method we applied it to a ring toss game, and generated robot motions for a physical humanoid robot. Styles from each of three random players were extracted automatically from their demonstrations, and used for generating robot motions. The robot imitates the styles of each player without exceeding the limitation of its physical constraints, while tossing the rings to the goal.
Takahiro Okamoto, Takaaki Shiratori, M. Glisson, K. Yamane, Shunsuke Kudoh, Katsushi Ikeuchi
IROS6
2014 Visibility-based blending for real-time applications
abstract
There are many situations in which virtual objects are presented half-transparently on a background in real time applications. In such cases, we often want to show the object with constant visibility. However, using the conventional alpha blending, visibility of a blended object substantially varies depending on colors, textures, and structures of the background scene. To overcome this problem, we present a framework for blending images based on a subjective metric of visibility. In our method, a blending parameter is locally and adaptively optimized so that visibility of each location achieves the targeted level. To predict visibility of an object blended by an arbitrary parameter, we utilize one of the error visibility metrics that have been developed for image quality assessment. In this study, we demonstrated that the metric we used can linearly predict visibility of a blended pattern on various texture images, and showed that the proposed blending methods can work in practical situations assuming augmented reality.
Taiki Fukiage, Takeshi Oishi, Katsushi Ikeuchi
ISMAR3
2014 Turbidity-based aerial perspective rendering for mixed reality
abstract
In outdoor Mixed Reality (MR), objects distant from the observer suffer from an effect called aerial perspective that fades the color of the objects and blends it to the environmental light color. The aerial perspective can be modeled using a physics-based approach; however, handling the changing and unpredictable environmental illumination is demanding. We present a turbidity-based method for rendering a virtual object with aerial perspective effect in a MR application. The proposed method first estimates the turbidity by matching luminance distributions of sky models and a captured omnidirectional sky image. Then the obtained turbidity is used to render the virtual object with aerial perspective.
Carlos Morales, Takeshi Oishi, Katsushi Ikeuchi
ISMAR3
2014 Bi-Polynomial Modeling of Low-Frequency Reflectances
abstract
We present a bi-polynomial reflectance model that can precisely represent the low-frequency component of reflectance. Most existing reflectance models aim at accurately representing the complete reflectance domain for photo-realistic rendering purposes. In contrast, our bi-polynomial model is developed for the purpose of accurately solving inverse problems by effectively discarding the high-frequency component while retaining nonlinear variations in the low-frequency part. The bi-polynomial reflectance model is useful for estimating reflectance and shape of an object. Experimental evaluation in comparison with other parametric reflectance models demonstrates that the proposed model achieves better performance in reflectometry and photometric stereo applications.
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Special Issue on ICPR 2012 Awarded Papers
Kim Boyer, Horst Bunke, Alberto Del Bimbo, Katsushi Ikeuchi, Ken-ichi Maeda
Pattern Recognit. Lett.4
2014 Toward a Dancing Robot With Listening Capability: Keypose-Based Integration of Lower-, Middle-, and Upper-Body Motions for Varying Music Tempos
abstract
This paper presents the development toward a dancing robot that can listen to and dance along with musical performances. One of the key components of this robot is the ability to modify its dance motions with varying tempos, without exceeding motor limitations, in the same way that human dancers modify their motions. In this paper, we first observe human performances with varying musical tempos of the same musical piece, and then analyze human modification strategies. The analysis is conducted in terms of three body components: lower, middle, and upper bodies. We assume that these body components have different purposes and different modification strategies, respectively, for the performance of a dance. For all of the motions of these three components, we have found that certain fixed postures, which we call keyposes, tend to be preserved. Thus, this paper presents a method to create motions for robots at a certain music tempo, from human motion at an original music tempo, by using these keyposes. We have implemented these algorithms as an automatic process and validated their effectiveness by using a physical humanoid robot HRP-2. This robot succeeded in performing the Aizu-bandaisan dance, one of the Japanese traditional folk dances, 1.2 and 1.5 times faster than the tempo originally learned, while maintaining its physical constraints. Although we are not achieving a dancing robot which autonomously interacts with varying music tempos, we think that our method has a vital role in the dancing-to-music capability.
Takahiro Okamoto, Takaaki Shiratori, Shunsuke Kudoh, Shinichiro Nakaoka, Katsushi Ikeuchi
IEEE Trans. Robotics5
2013 Adherent Raindrop Detection and Removal in Video
abstract
Raindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. First, it detects raindrops based on the motion and the intensity temporal derivatives of the input video. Second, relying on an analysis that some areas of a raindrop completely occludes the scene, yet the remaining areas occludes only partially, the method removes the two types of areas separately. For partially occluding areas, it restores them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity change. For completely occluding areas, it recovers them by using a video completion technique. Experimental results using various real videos show the effectiveness of the proposed method.
Shaodi You, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi
CVPR4
2013 Beyond Point Clouds: Scene Understanding by Reasoning Geometry and Physics
abstract
In this paper, we present an approach for scene understanding by reasoning physical stability of objects from point cloud. We utilize a simple observation that, by human design, objects in static scenes should be stable with respect to gravity. This assumption is applicable to all scene categories and poses useful constraints for the plausible interpretations (parses) in scene understanding. Our method consists of two major steps: 1) geometric reasoning: recovering solid 3D volumetric primitives from defective point cloud, and 2) physical reasoning: grouping the unstable primitives to physically stable objects by optimizing the stability and the scene prior. We propose to use a novel disconnectivity graph (DG) to represent the energy landscape and use a Swendsen-Wang Cut (MCMC) method for optimization. In experiments, we demonstrate that the algorithm achieves substantially better performance for i) object segmentation, ii) 3D volumetric recovery of the scene, and iii) better parsing result for scene understanding in comparison to state-of-the-art methods in both public dataset and our own new dataset.
Bo Zheng 0001, Yibiao Zhao, Joey C. Yu, Katsushi Ikeuchi, Song-Chun Zhu
CVPR4
2013 Representation and mapping of dexterous manipulation through task primitives
abstract
The goal of this work is to teach a robot to regrasp an object using knowledge obtained from human demonstration. This paper presents a task model that represents a human regrasping movement. The task model is based on the topological information and comprised of four task primitives. Human regrasping movement is recognised and represented as a sequence of these task primitives by the proposed recognition algorithm. The proposed method then maps each task primitive to the target robot hand using knowledge obtained from human demonstration. The experimental result verified the proposed task model by executing the regrasping movement on the real robot hand.
Phongtharin Vinayavekhin, Shunsuke Kudoh, Jun Takamatsu, Yoshihiro Sato, Katsushi Ikeuchi
ICRA5
2013 A coarse-to-fine IP-driven registration for pose estimation from single ultrasound image
Bo Zheng 0001, Ryoichi Ishikawa, Jun Takamatsu, Takeshi Oishi, Katsushi Ikeuchi
Comput. Vis. Image Underst.5
2013 Camera Spectral Sensitivity and White Balance Estimation from Sky Images
abstract
Photometric camera calibration is often required in physics-based computer vision. There have been a number of studies to estimate camera response functions (gamma function), and vignetting effect from images. However less attention has been paid to camera spectral sensitivities and white balance settings. This is unfortunate, since those two properties significantly affect image colors. Motivated by this, a method to estimate camera spectral sensitivities and white balance setting jointly from images with sky regions is introduced. The basic idea is to use the sky regions to infer the sky spectra. Given sky images as the input and assuming the sun direction with respect to the camera viewing direction can be extracted, the proposed method estimates the turbidity of the sky by fitting the image intensities to a sky model. Subsequently, it calculates the sky spectra from the estimated turbidity. Having the sky $$RGB$$ values and their corresponding spectra, the method estimates the camera spectral sensitivities together with the white balance setting. Precomputed basis functions of camera spectral sensitivities are used in the method for robust estimation. The whole method is novel and practical since, unlike existing methods, it uses sky images without additional hardware, assuming the geolocation of the captured sky is known. Experimental results using various real images show the effectiveness of the method.
Rei Kawakami, Hongxun Zhao, Robby T. Tan, Katsushi Ikeuchi
Int. J. Comput. Vis.4
2013 A Branch-and-Bound Approach to Correspondence and Grouping Problems
abstract
Data correspondence/grouping under an unknown parametric model is a fundamental topic in computer vision. Finding feature correspondences between two images is probably the most popular application of this research field, and is the main motivation of our work. It is a key ingredient for a wide range of vision tasks, including three-dimensional reconstruction and object recognition. Existing feature correspondence methods are based on either local appearance similarity or global geometric consistency or a combination of both in some heuristic manner. None of these methods is fully satisfactory, especially in the presence of repetitive image textures or mismatches. In this paper, we present a new algorithm that combines the benefits of both appearance-based and geometry-based methods and mathematically guarantees a global optimization. Our algorithm accepts the two sets of features extracted from two images as input, and outputs the feature correspondences with the largest number of inliers, which verify both the appearance similarity and geometric constraints. Specifically, we formulate the problem as a mixed integer program and solve it efficiently by a series of linear programs via a branch-and-bound procedure. We subsequently generalize our framework in the context of data correspondence/grouping under an unknown parametric model and show it can be applied to certain classes of computer vision problems. Our algorithm has been validated successfully on synthesized data and challenging real images.
Jean-Charles Bazin, Hongdong Li, In-So Kweon, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 Radiometric Calibration by Rank Minimization
abstract
We present a robust radiometric calibration framework that capitalizes on the transform invariant low-rank structure in the various types of observations, such as sensor irradiances recorded from a static scene with different exposure times, or linear structure of irradiance color mixtures around edges. We show that various radiometric calibration problems can be treated in a principled framework that uses a rank minimization approach. This framework provides a principled way of solving radiometric calibration problems in various settings. The proposed approach is evaluated using both simulation and real-world datasets and shows superior performance to previous approaches.
Joon-Young Lee, Yasuyuki Matsushita, Boxin Shi, In-So Kweon, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 Globally optimal line clustering and vanishing point estimation in Manhattan world
abstract
The projections of world parallel lines in an image intersect at a single point called the vanishing point (VP). VPs are a key ingredient for various vision tasks including rotation estimation and 3D reconstruction. Urban environments generally exhibit some dominant orthogonal VPs. Given a set of lines extracted from a calibrated image, this paper aims to (1) determine the line clustering, i.e. find which line belongs to which VP, and (2) estimate the associated orthogonal VPs. None of the existing methods is fully satisfactory because of the inherent difficulties of the problem, such as the local minima and the chicken-and-egg aspect. In this paper, we present a new algorithm that solves the problem in a mathematically guaranteed globally optimal manner and can inherently enforce the VP orthogonality. Specifically, we formulate the task as a consensus set maximization problem over the rotation search space, and further solve it efficiently by a branch-and-bound procedure based on the Interval Analysis theory. Our algorithm has been validated successfully on sets of challenging real images as well as synthetic data sets.
Jean-Charles Bazin, Yongduek Seo, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi, In-So Kweon, Marc Pollefeys
CVPR5
2012 A biquadratic reflectance model for radiometric image analysis
abstract
Radiometric image analysis methods heavily rely on reflectance models. Due to the complexity of real materials, methods based on simple models such as the Lambertian model often suffer from inaccuracy. On the other hand, more advanced models such as the Cook-Torrance model severely complicate the analysis problem. We tackle this dilemma by focusing on the low-frequency component of the reflectance. We propose a compact biquadratic reflectance model to represent the reflectance of a broad class of materials precisely in the low-frequency domain. We validate our model by fitting to both existing parametric models and non-parametric measured data, and show that our model outperforms existing parametric diffuse models. We show applications of reflectometry using general diffuse surfaces and photometric stereo for general isotropic materials. Experimental results show the effectiveness of our biquadratic model and its usefulness in radiometric image analysis.
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
CVPR4
2012 Elevation Angle from Reflectance Monotonicity: Photometric Stereo for General Isotropic Reflectances
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
ECCV (3)4
2012 Reduction of contradictory partial occlusion in mixed reality by using characteristics of transparency perception
abstract
One of the challenges in mixed reality (MR) applications is handling contradictory occlusions between real and virtual objects. The previous studies have tried to solve the occlusion problem by extracting the foreground region from the real image. However, real-time occlusion handling is still difficult since it takes too much computational cost to precisely segment foreground regions in a complex scene. In this study, therefore, we proposed an alternative solution to the occlusion problem that does not require precise foreground-background segmentation. In our method, a virtual object is blended with a real scene so that the virtual object can be perceived as being behind the foreground region. For this purpose, we first investigated characteristics of human transparency perception in a psychophysical experiment. Then we made a blending algorithm applicable to real scenes based on the results of the experiment.
Taiki Fukiage, Takeshi Oishi, Katsushi Ikeuchi
ISMAR3
2012 Achieving robust alignment for outdoor mixed reality using 3D range data
abstract
Mixed reality (MR) technology can be applied to various applications such as architecture, advertising, and navigation systems, so the desire to utilize MR in outdoor environments has been increasing. In order to utilize MR, it is necessary to achieve alignment super imposing virtual contents in the desired position. However, because light changes continually in outdoor environments, and the appearance of real objects changes also, in some cases the previous image-based alignment methods do not work well. In this paper, a robust image-based alignment method to be used in outdoor environments is proposed. In the proposed method, the albedo of real objects is estimated using 3D shapes of these objects in advance, and the appearance is reproduced from the albedo and current light environment. The appearance of real objects and reproduced image becomes close, so a robust image-based alignment is achieved.
Masaki Inaba, Atsuhiko Banno, Takeshi Oishi, Katsushi Ikeuchi
VRST4
2012 Estimation of F-Matrix and image rectification by double quaternion
Atsuhiko Banno, Katsushi Ikeuchi
Inf. Sci.2
2011 High-resolution hyperspectral imaging via matrix factorization
abstract
Hyperspectral imaging is a promising tool for applications in geosensing, cultural heritage and beyond. However, compared to current RGB cameras, existing hyperspectral cameras are severely limited in spatial resolution. In this paper, we introduce a simple new technique for reconstructing a very high-resolution hyperspectral image from two readily obtained measurements: A lower-resolution hyper-spectral image and a high-resolution RGB image. Our approach is divided into two stages: We first apply an unmixing algorithm to the hyperspectral input, to estimate a basis representing reflectance spectra. We then use this representation in conjunction with the RGB input to produce the desired result. Our approach to unmixing is motivated by the spatial sparsity of the hyperspectral input, and casts the unmixing problem as the search for a factorization of the input into a basis and a set of maximally sparse coefficients. Experiments show that this simple approach performs reasonably well on both simulations and real data examples.
Rei Kawakami, Yasuyuki Matsushita, John Wright 0001, Moshe Ben-Ezra, Yu-Wing Tai, Katsushi Ikeuchi
CVPR6
2011 Radiometric calibration by transform invariant low-rank structure
abstract
We present a robust radiometric calibration method that capitalizes on the transform invariant low-rank structure of sensor irradiances recorded from a static scene with different exposure times. We formulate the radiometric calibration problem as a rank minimization problem. Unlike previous approaches, our method naturally avoids over-fitting problem; therefore, it is robust against biased distribution of the input data, which is common in practice. When the exposure times are completely unknown, the proposed method can robustly estimate the response function up to an exponential ambiguity. The method is evaluated using both simulation and real-world datasets and shows a superior performance than previous approaches.
Joon-Young Lee, Boxin Shi, Yasuyuki Matsushita, In-So Kweon, Katsushi Ikeuchi
CVPR5
2011 Locally rigid globally non-rigid surface registration
abstract
We present a novel non-rigid surface registration method that achieves high accuracy and matches characteristic features without manual intervention. The key insight is to consider the entire shape as a collection of local structures that individually undergo rigid transformations to collectively deform the global structure. We realize this locally rigid but globally non-rigid surface registration with a newly derived dual-grid Free-form Deformation (FFD) framework. We first represent the source and target shapes with their signed distance fields (SDF). We then superimpose a sampling grid onto a conventional FFD grid that is dual to the control points. Each control point is then iteratively translated by a rigid transformation that minimizes the difference between two SDFs within the corresponding sampling region. The translated control points then interpolate the embedding space within the FFD grid and determine the overall deformation. The experimental results clearly demonstrate that our method is capable of overcoming the difficulty of preserving and matching local features.
Kent Fujiwara, Ko Nishino, Jun Takamatsu, Bo Zheng 0001, Katsushi Ikeuchi
ICCV5
2011 Towards an automatic robot regrasping movement based on human demonstration using tangle topology
abstract
This paper introduces a novel method to teach a robot to regrasp an object based on the Programming by Demonstration paradigm. In this paradigm, a robot observes a human performing a regrasping task via various sensors to recognise crucial information in order to reproduce the task using its own hand. The main contribution is in the proposal of a representation technique that can analyse a human regrasping movement and reproduce this movement in a robot hand. The technique is based on tangle topology where both hand and manipulated object are considered as strands. This allows a regrasping movement to be considered as an alteration of the tangle relationship between the strands (hand and the object) over time. Human regrasping movements are analysed and reproduced in multi-fingered robot hands in a grasp simulation to demonstrate the efficiency of the proposed method.
Phongtharin Vinayavekhin, Shunsuke Kudoh, Katsushi Ikeuchi
ICRA3
2011 Disparity map refinement and 3D surface smoothing via directed anisotropic diffusion
Atsuhiko Banno, Katsushi Ikeuchi
Comput. Vis. Image Underst.2
2011 Image-Based Network Rendering of Large Meshes for Cloud Computing
Yasuhide Okamoto, Takeshi Oishi, Katsushi Ikeuchi
Int. J. Comput. Vis.3
2011 Determination of motion parameters of a moving range sensor approximated by polynomials for rectification of distorted 3D data
Atsuhiko Banno, Katsushi Ikeuchi
Mach. Vis. Appl.2
2010 Consensus photometric stereo
abstract
This paper describes a photometric stereo method that works with a wide range of surface reflectances. Unlike previous approaches that assume simple parametric models such as Lambertian reflectance, the only assumption that we make is that the reflectance has three properties; monotonicity, visibility, and isotropy with respect to the cosine of light direction and surface orientation. In fact, these properties are observed in many non-Lambertian diffuse reflectances. We also show that the monotonicity and isotropy properties hold specular lobes with respect to the cosine of the surface orientation and the bisector between the light direction and view direction. Each of these three properties independently gives a possible solution space of the surface orientation. By taking the intersection of the solution spaces, our method determines the surface orientation in a consensus manner. Our method naturally avoids the need for radiometrically calibrating cameras because the radiometric response function preserves these three properties. The effectiveness of the proposed method is demonstrated using various simulated and real-world scenes that contain a variety of diffuse and specular surfaces.
Tomoaki Higo, Yasuyuki Matsushita, Katsushi Ikeuchi
CVPR3
2010 Estimating optical properties of layered surfaces using the spider model
abstract
Many object surfaces are composed of layers of different physical substances, known as layered surfaces. These surfaces, such as patinas, water colors, and wall paintings, have more complex optical properties than diffuse surfaces. Although the characteristics of layered surfaces, like layer opacity, mixture of colors, and color gradations, are significant, they are usually ignored in the analysis of many methods in computer vision, causing inaccurate or even erroneous results. Therefore, the main goals of this paper are twofold: to solve problems of layered surfaces by focusing mainly on surfaces with two layers (i.e., top and bottom layers), and to introduce a decomposition method based on a novel representation of a nonlinear correlation in the color space that we call the “spider” model. When we plot a mixture of colors of one bottom layer and n different top layers into the RGB color space, then we will have n different curves intersecting at one point, resembling the shape of a spider. Hence, given a single input image containing one bottom layer and at least one top layer, we can fit their color distributions by using the spider model and then decompose those layered surfaces. The last step is equivalent to extracting the approximated optical properties of the two layers: the top layer's opacity, and the top and bottom layers' reflections. Experiments with real images, which include the photographs of ancient wall paintings, show the effectiveness of our method.
Tetsuro Morimoto, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi
CVPR4
2010 Estimating demosaicing algorithms using image noise variance
abstract
We propose a method for estimating demosaicing algorithms from image noise variance. We show that the noise variance in interpolated pixels becomes smaller than that of directly observed pixels without interpolation. Our method capitalizes on the spatial variation of image noise variance in demosaiced images to estimate the color filter array patterns and demosaicing algorithms. We verify the effectiveness of the proposed method using various images demosaiced with different demosaicing algorithms extensively.
Jun Takamatsu, Yasuyuki Matsushita, Tsukasa Ogasawara, Katsushi Ikeuchi
CVPR4
2010 Photometric stereo under unknown light sources using robust SVD with missing data
abstract
In this paper, we propose a novel photometric stereo method that uses singular value decomposition. Singular value decomposition can solve the photometric stereo problem when the light source direction is unknown; however, it has the critical problem of being sensitive to outliers. We therefore propose a novel singular value decomposition method that is robust to outliers. We also show some results of our photometric stereo method when applied to objects that involve not only diffuse reflection but also specular reflection.
Daisuke Miyazaki, Katsushi Ikeuchi
ICIP2
2010 Temporal scaling of leg motion for music feedback system of a dancing humanoid robot
abstract
In this paper, we propose a method to achieve temporal scaling of leg motions as a fundamental technique for a music feedback system of a dancing humanoid robot. We asked dancers to perform dance motion at normal musical tempo and faster musical tempos and observed how dancers modified performance for given musical tempos. The obtained insights from the observation are 1) a dancer needs to preserve leg postures that are important to emphasize dance expression, 2) there is a priority to determine what features of leg motion can be adjusted, and 3) stylistic leg motion resembles normal step motion if dancers cannot follow fast musical tempo completely. Based on these insights, we generate leg motion appropriately adjusted for changing musical tempo while maintaining balance. We validated our method via simulation experiments with a humanoid robot HRP-2.
Takahiro Okamoto, Takaaki Shiratori, Shunsuke Kudoh, Katsushi Ikeuchi
IROS4
2010 Detecting dance motion structure using body components and turning motions
abstract
This paper presents a novel method for robust dance motion structure detection. In the japanese folk dance domain, teachers created illustrations of dance poses. These poses characterize the most important movements of a dance. So far there is no simple and reliable extraction method which can extract all poses as shown in these drawings. We use these poses for the Task Model (TM) in the context of Learning from Observation (LFO). LFO which is a well known technique for successful human to robot motion mapping, consists of tasks (what to do) and skills (how to do). We propose a novel approach, to extract special motions from a dance, called turning motions useful for skill mapping in the LFO paradigm. Furthermore, we use a modified version of this approach, to detect all poses as shown in the drawings, called turning poses. To achieve this we observe both forearms at the same time and analyze their movement in different 2-D coordinate planes. We evaluate the parameters with and without a weighting function where we minimize acceleration, velocity and power. We successfully demonstrate this novel method using two very different japanese folk dances and discuss further implications of this work in respect to the LFO paradigm and dances of other domains.
Bjoern Rennhak, Takaaki Shiratori, Shunsuke Kudoh, Phongtharin Vinayavekhin, Katsushi Ikeuchi
IROS5
2010 Foreground and shadow occlusion handling for outdoor augmented reality
abstract
Occlusion handling in augmented reality (AR) applications is challenging in synthesizing virtual objects correctly into the real scene with respect to existing foregrounds and shadows. Furthermore, outdoor environment makes the task more difficult due to the unpredictable illumination changes. This paper proposes novel outdoor illumination constraints for resolving the foreground occlusion problem in outdoor environment. The constraints can be also integrated into a probabilistic model of multiple cues for a better segmentation of the foreground. In addition, we introduce an effective method to resolve the shadow occlusion problem by using shadow detection and recasting with a spherical vision camera. We have applied the system in our digital cultural heritage project named Virtual Asuka (VA) and verified the effectiveness of the system.
Boun Vinh Lu, Tetsuya Kakuta, Rei Kawakami, Takeshi Oishi, Katsushi Ikeuchi
ISMAR5
2010 Omnidirectional texturing based on robust 3D registration through Euclidean reconstruction from two spherical images
Atsuhiko Banno, Katsushi Ikeuchi
Comput. Vis. Image Underst.2
2010 Editorial for the Special Issue on Photometric Analysis for Computer Vision
Peter N. Belhumeur, Katsushi Ikeuchi, Emmanuel Prados, Stefano Soatto, Peter F. Sturm
Int. J. Comput. Vis.2
2010 Median Photometric Stereo as Applied to the Segonko Tumulus and Museum Objects
Daisuke Miyazaki, Kenji Hara, Katsushi Ikeuchi
Int. J. Comput. Vis.3
2010 Wavelet-Texture Method: Appearance Compression by Polarization, Parametric Reflection Model, and Daubechies Wavelet
Daisuke Miyazaki, Takushi Shibata, Katsushi Ikeuchi
Int. J. Comput. Vis.3
2010 An Adaptive and Stable Method for Fitting Implicit Polynomial Curves and Surfaces
abstract
Representing 2D and 3D data sets with implicit polynomials (IPs) has been attractive because of its applicability to various computer vision issues. Therefore, many IP fitting methods have already been proposed. However, the existing fitting methods can be and need to be improved with respect to computational cost for deciding on the appropriate degree of the IP representation and to fitting accuracy, while still maintaining the stability of the fit. We propose a stable method for accurate fitting that automatically determines the moderate degree required. Our method increases the degree of IP until a satisfactory fitting result is obtained. The incrementability of QR decomposition with Gram-Schmidt orthogonalization gives our method computational efficiency. Furthermore, since the decomposition detects the instability element precisely, our method can selectively apply ridge regression-based constraints to that element only. As a result, our method achieves computational stability while maintaining fitting accuracy. Experimental results demonstrate the effectiveness of our method compared with prior methods.
Bo Zheng 0001, Jun Takamatsu, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Interactive Shadow Removal from a Single Image Using Hierarchical Graph Cut
Daisuke Miyazaki, Yasuyuki Matsushita, Katsushi Ikeuchi
ACCV (1)3
2009 Multilevel Algebraic Invariants Extraction by Incremental Fitting Scheme
Bo Zheng 0001, Jun Takamatsu, Katsushi Ikeuchi
ACCV (1)3
2009 Color estimation from a single surface color
abstract
This paper estimates illumination colors by using only a single surface color taken under multiple illumination colors. Past researchers have found that there is a difficulty in estimating illumination colors using a single surface color. However, the method presented here overcomes the problem. Surface color is estimated by considering four characteristics of illumination and surface color spaces. First, the outdoor-illumination colors exist in a specific color range. Second, multiple illuminations give constraints for the surface color. Third, multiple illuminations also give constraints for the color range. Fourth, each color component affects those constraints in a different manner. Based on those characteristics, a novel method can be designed. The proposed method produces consistently accurate results when multiple illumination colors are used, because the constraint (possible range of illumination colors) on illumination colors refines the estimated illumination colors, effectively.
Rei Kawakami, Katsushi Ikeuchi
CVPR2
2009 A hand-held photometric stereo camera for 3-D modeling
abstract
This paper presents a simple yet practical 3-D modeling method for recovering surface shape and reflectance from a set of images. We attach a point light source to a hand-held camera to add a photometric constraint to the multi-view stereo problem. Using the photometric constraint, we simultaneously solve for shape, surface normal, and reflectance. Unlike prior approaches, we formulate the problem using realistic assumptions of a near light source, non-Lambertian surfaces, perspective camera model, and the presence of ambient lighting. The effectiveness of the proposed method is verified using simulated and real-world scenes.
Tomoaki Higo, Yasuyuki Matsushita, Neel Joshi, Katsushi Ikeuchi
ICCV4
2008 Estimating camera response functions using probabilistic intensity similarity
abstract
We propose a method for estimating camera response functions using a probabilistic intensity similarity measure. The similarity measure represents the likelihood of two intensity observations corresponding to the same scene radiance in the presence of noise. We show that the response function and the intensity similarity measure are strongly related. Our method requires several input images of a static scene taken from the same viewing position with fixed camera parameters. Noise causes pixel values at the same pixel coordinate to vary in these images, even though they measure the same scene radiance. We use these fluctuations to estimate the response function by maximizing the intensity similarity function for all pixels. Unlike prior noise-based estimation methods, our method requires only a small number of images, so it works with digital cameras as well as video cameras. Moreover, our method does not rely on any special image processing or statistical prior models. Real-world experiments using different cameras demonstrate the effectiveness of the technique.
Jun Takamatsu, Yasuyuki Matsushita, Katsushi Ikeuchi
CVPR3
2008 Estimating Radiometric Response Functions from Image Noise Variance
Jun Takamatsu, Yasuyuki Matsushita, Katsushi Ikeuchi
ECCV (4)3
2008 Integrating region growing and classification for segmentation and matting
abstract
This paper presents a supervised foreground segmentation method that uses local and global feature similarity with edge constraint. This framework integrates and extends the notion of region growing and classification to deal with local and global fitness. It parameterizes constraint of growing using Chebyshev's inequality. The constraint is used to stop segmentation before matting. Matting relies on both local and global information. The proposed method outperforms many of the current methods in the sense of correctness and minimal user interaction, and it does so in a reasonable computation time.
Miti Ruchanurucks, Koichi Ogawara, Katsushi Ikeuchi
ICIP3
2008 Real-time image-based rendering system for virtual city based on image compression technique and eigen texture method
abstract
Computer modeling of a large-scale scene such as a city becomes an important topic for computer vision and computer graphics research areas etc. Image-based rendering (IBR) is an effective method for expressing realistic scene, and can construct any arbitrary viewpoint by using the captured real images. However, the large size of the image database in IBR causes serious problems in actual applications, leading to the use of compression techniques. We propose a compression technique based on eigen space combined with a block matching technique to get better result. We also propose a technique to restore the compressed data on Graphic Processing Unit (GPU), allowing us to perform high-speed rendering without raising the load on the CPU.
Ryo Sato, Shintaro Ono, Hiroshi Kawasaki, Katsushi Ikeuchi
ICPR4
2008 Detection of moving objects and cast shadows using a spherical vision camera for outdoor mixed reality
abstract
This paper presents a method to detect moving objects and remove their shadows for superimposing them on Mixed Reality (MR) systems. We cut out the foreground from a real image using a probability-based segmentation method. Using color, spatial, and temporal priors, we can improve the accuracy of the segmentation. Energy minimization is executed by graph cuts. Then we remove the shadow region from the foreground with F-value calculated from the pixel value and the spectral sensitivity characteristic of the camera. Finally we superimpose virtual objects using the stencil buffer, which is used to limit the area of rendering for each pixel. Synthesized images of an outdoor scene show the efficiency of the proposed method.
Tetsuya Kakuta, Boun Vinh Lu, Rei Kawakami, Takeshi Oishi, Katsushi Ikeuchi
VRST5
2008 Flying Laser Range Sensor for Large-Scale Site-Modeling and Its Applications in Bayon Digital Archival Project
Atsuhiko Banno, Tomohito Masuda, Takeshi Oishi, Katsushi Ikeuchi
Int. J. Comput. Vis.4
2008 Mixture of Spherical Distributions for Single-View Relighting
abstract
We present a method for simultaneously estimating the illumination of a scene and the reflectance property of an object from single view images - a single image or a small number of images taken from the same viewpoint. We assume that the illumination consists of multiple point light sources and the shape of the object is known. First, we represent the illumination on the surface of a unit sphere as a finite mixture of von Mises-Fisher distributions based on a novel spherical specular reflection model that well approximates the Torrance-Sparrow reflection model. Next, we estimate the parameters of this mixture model including the number of its component distributions and the standard deviation of them, which correspond to the number of light sources and the surface roughness, respectively. Finally, using these results as the initial estimates, we iteratively refine the estimates based on the original Torrance-Sparrow reflection model. The final estimates can be used to relight single-view images such as altering the intensities and directions of the individual light sources. The proposed method provides a unified framework based on directional statistics for simultaneously estimating the intensities and directions of an unknown number of light sources as well as the specular reflection parameter of the object in the scene.
Kenji Hara, Ko Nishino, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Hole Filling of a 3D Model by Flipping Signs of a Signed Distance Field in Adaptive Resolution
abstract
When we use range finders to observe the shape of an object, many occluded areas may occur. These become holes and gaps in the model and make it undesirable for various applications. We propose a novel method to fill holes and gaps to complete this incomplete model. As an intermediate representation, we use a Signed Distance Field (SDF), which stores Euclidean signed distances from a voxel to the nearest point of the mesh model. By using an SDF, we can obtain interpolating surfaces for holes and gaps. The proposed method generates an interpolating surface that becomes smoothly continuous with real surfaces by minimizing the area of the interpolating surface. Since the isosurface of an SDF can be identified as being a real or interpolating surface from the magnitude of signed distances, our method computes the area of an interpolating surface in the neighborhood of a voxel both before and after flipping the sign of the signed distance of the voxel. If the area is reduced by flipping the sign, our method changes the sign for the voxel. Therefore, we minimize the area of the interpolating surface by iterating this computation until convergence. Unlike methods based on Partial Differential Equations (PDE), our method does not require any boundary condition, and the initial state that we use is automatically obtained by computing the distance to the closest point of the real surface. Moreover, because our method can be applied to an SDF of adaptive resolution, our method efficiently interpolates large holes and gaps of high curvature. We tested the proposed method with both synthesized and real objects and evaluated the interpolating surfaces.
Ryusuke Sagawa, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Adaptively Determining Degrees of Implicit Polynomial Curves and Surfaces
Bo Zheng 0001, Jun Takamatsu, Katsushi Ikeuchi
ACCV (2)3
2007 Marker-less Human Motion Estimation using Articulated Deformable Model
abstract
This paper presents a novel whole body motion estimation method by fitting a deformable articulated model of the human body into the 3D reconstructed volume obtained from multiple video streams. The advantage of the proposed method is two fold: (1) combination of a robust estimator and ICP algorithm with Kd-tree search in pose and normal space make it possible to track complex and dynamic motion robustly against noise and interference between limb and torso, (2) the hierarchical estimation and backtrack re-estimation algorithm enable accurate estimation. The power to track challenging whole body motion in real environment is also presented.
Koichi Ogawara, Katsushi Ikeuchi
ICRA3
2007 Humanoid Robot Painter: Visual Perception and High-Level Planning
abstract
This paper presents visual perception discovered in high-level manipulator planning for a robot to reproduce the procedure involved in human painting. First, we apply a technique of 2D object segmentation that considers region similarity as an objective function and edge as a constraint with artificial intelligent used as a criterion function. The system can segment images more effectively than most of existing methods, even if the foreground is very similar to the background. Second, we propose a novel color perception model that shows similarity to human perception. The method outperforms many existing color reduction algorithms. Third, we propose a novel global orientation map perception using a radial basis function. Finally, we use the derived model along with the brush's position- and force-sensing to produce a visual feedback drawing. Experiments show that our system can generate good paintings including portraits.
Miti Ruchanurucks, Shunsuke Kudoh, Koichi Ogawara, Takaaki Shiratori, Katsushi Ikeuchi
ICRA5
2007 Multilinear analysis for task recognition and person identification
abstract
This paper introduces a Multi Factor Tensor(MFT) model to recognize motion styles and person identities in dance sequences. We apply a musical information analysis method in segmenting the motion sequence relevant to the key poses and the musical rhythm. We define a task model considering the repeated motion segments, where the motion is decomposed into person invariant factor task and person dependant factor style. We capture the motion data of different people for a few cycles, segment it using the musical analysis approach, normalize the segments using a vectorization method, and realize our MFT model. The experiments are conducted according to two approaches. Various experiments that we conduct to evaluate the potential of the recognition ability of our proposed approaches and the results demonstrate the high accuracy of our model. The recognition results and the motion decomposition will be used in further extending the motion generation process in various styles and for different tasks.
Manoj Perera, Takaaki Shiratori, Shunsuke Kudoh, Atsushi Nakazawa, Katsushi Ikeuchi
IROS5
2007 Robot painter: from object to trajectory
abstract
This paper presents visual perception discovered in high-level manipulator planning for a robot to reproduce the procedure involved in human painting. First, we propose a technique of 3D object segmentation that can work well even when the precision of the cameras is inadequate. Second, we apply a simple yet powerful fast color perception model that shows similarity to human perception. The method outperforms many existing interactive color perception algorithms. Third, we generate global orientation map perception using a radial basis function. Finally, we use the derived foreground, color segments, and orientation map to produce a visual feedback drawing. Our main contributions are 3D object segmentation and color perception schemes.
Miti Ruchanurucks, Shunsuke Kudoh, Koichi Ogawara, Takaaki Shiratori, Katsushi Ikeuchi
IROS5
2007 Temporal scaling of upper body motion for Sound feedback system of a dancing humanoid robot
abstract
This paper proposes a method to model the modification of upper body motion of dance performance based on the speed of played music. When we observed structured dance motion performed at a normal music playback speed and motion performed at a faster music playback speed, we found that the detail of each motion is slightly different while the whole of the dance motion is similar in both cases. This phenomenon is derived from the fact that dancers omit the details and perform the essential part of the dance in order to follow the faster speed of the music. To clarify this phenomenon, we analyzed the motion differences in the frequency domain, and obtained two insights on the omission of motion details: (1) High frequency components are gradually attenuated depending on the musical speed, and (2) important stop motions are preserved even when high frequency components are attenuated. Based on these insights, we modeled our motion modification considering musical speed and joint limitations that a humanoid robot has. We show the effectiveness of our method via some applications for humanoid robot motion generation.
Takaaki Shiratori, Shunsuke Kudoh, Shinichiro Nakaoka, Katsushi Ikeuchi
IROS4
2007 Editorial
Katsushi Ikeuchi, Gudrun Klinker, Yuichi Ohta, Richard Szeliski
Int. J. Comput. Vis.1
2007 The Great Buddha Project: Digitally Archiving, Restoring, and Analyzing Cultural Heritage Objects
Katsushi Ikeuchi, Takeshi Oishi, Jun Takamatsu, Ryusuke Sagawa, Atsushi Nakazawa, Ryo Kurazume, Ko Nishino, Mawo Kamakura, Yasuhide Okamoto
Int. J. Comput. Vis.1
2007 The separation of reflected and transparent layers from real-world image sequence
Thanda Oo, Hiroshi Kawasaki, Yutaka Ohsawa, Katsushi Ikeuchi
Mach. Vis. Appl.4
2007 Shape Estimation of Transparent Objects by Using Inverse Polarization Ray Tracing
abstract
Few methods have been proposed to measure three-dimensional shapes of transparent objects such as those made of glass and acrylic. In this paper, we propose a novel method for estimating the surface shapes of transparent objects by analyzing the polarization state of the light. Existing methods do not fully consider the reflection, refraction, and transmission of the light occurring inside a transparent object. We employ a polarization raytracing method to compute both the path of the light and its polarization state. Polarization raytracing is a combination of conventional raytracing, which calculates the trajectory of light rays, and Mueller calculus, which calculates the polarization state of the light. First, we set an initial value of the shape of the transparent object. Then, by changing the shape, the method minimizes the difference between the input polarization data and the rendered polarization data calculated by polarization raytracing. Finally, after the iterative computation is converged, the shape of the object is obtained. We also evaluate the method by measuring some real transparent objects.
Daisuke Miyazaki, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Separation of Reflection and Transparency Using Epipolar Plane Image Analysis
Thanda Oo, Hiroshi Kawasaki, Yutaka Ohsawa, Katsushi Ikeuchi
ACCV (1)4
2006 Stepping Motion for a Human-like Character to Maintain Balance against Large Perturbations
abstract
We propose a method of maintaining balance for a human-like character against large perturbations. The method enables a human-like model to maintain its balance with active whole-body motion, such as rotating its arms, bending down, and taking a step, if necessary. First, we capture the human motions of maintaining balance and abstract essential mechanisms from these motions. Next, we construct a model of maintaining balance that has a simple structure, such as an inverted pendulum. This model has two modes of maintaining balance: keeping the feet on the ground, and stepping. In this paper, the stepping mode is mainly described. Finally, we generate whole-body motion based on the model against several perturbations, and we discuss the validity of our method
Shunsuke Kudoh, Taku Komura, Katsushi Ikeuchi
ICRA3
2006 Humanoid Robot Motion Generation with Sequential Physical Constraints
abstract
This paper presents a method to optimize and filter trajectories generated from recorded human motion for a humanoid robot with physical limits. The objective function is responsible for mimicking human trainers, enhancing the possibility for fast convergence, while constraints are used to transform motion within the limit of the capabilities of the humanoid robot. Those constraints for angle, velocity, and dynamic force are represented as B-spline coefficients. An iterative soft-constraint paradigm is proposed to enhance the quality of velocity and force constraints. Collision avoidance is also considered as a constraint. For precision refinement, in regions of high-frequency motion not adequately modeled by an initial splining, B-spline is extensible into a hierarchy so that optimization that meets global criteria can be performed locally. Furthermore, all of the constraints can be used solely as filtering. To use these filters, an effective method to directly decompose a trajectory to a B-spline is also presented
Miti Ruchanurucks, Shinichiro Nakaoka, Shunsuke Kudoh, Katsushi Ikeuchi
ICRA4
2006 Synthesizing Dance Performance using Musical and Motion Features
abstract
This paper proposes a method for synthesizing dance performance synchronized to played music and our method presents a system that imitates dancers' skills in performing their motion while they listen to the music. Our method consists of a motion analysis, a music analysis, and a motion synthesis based on results of the analyses. In these analysis steps, motion and music features are acquired. These features are derived from motion keyframes, motion intensity, music intensity, musical beats, and chord changes. Our system also constructs a motion graph to search similar poses from given dance sequences and to connect them as possible transitions. In the synthesis step, the trajectory that provides the best correlation between music and motion features is selected from the motion graph, and the resulting motion is generated. Our experimental results indicate that our proposed method actually creates dance as the system "hears" the music
Takaaki Shiratori, Atsushi Nakazawa, Katsushi Ikeuchi
ICRA3
2006 Dancing-to-Music Character Animation
abstract
Abstract In computer graphics, considerable research has been conducted on realistic human motion synthesis. However, most research does not consider human emotional aspects, which often strongly affect human motion. This paper presents a new approach for synthesizing dance performance matched to input music, based on the emotional aspects of dance performance. Our method consists of a motion analysis, a music analysis, and a motion synthesis based on the extracted features. In the analysis steps, motion and music feature vectors are acquired. Motion vectors are derived from motion rhythm and intensity, while music vectors are derived from musical rhythm, structure, and intensity. For synthesizing dance performance, we first find candidate motion segments whose rhythm features are matched to those of each music segment, and then we find the motion segment set whose intensity is similar to that of music segments. Additionally, our system supports having animators control the synthesis process by assigning desired motion segments to the specified music segments. The experimental results indicate that our method actually creates dance performance as if a character was listening and expressively dancing to the music. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism Animation; J.5 [Arts and Humanities]: Performing Arts Music
Takaaki Shiratori, Atsushi Nakazawa, Katsushi Ikeuchi
Comput. Graph. Forum3
2006 Representation for knot-tying tasks
abstract
The learning from observation (LFO) paradigm has been widely applied in various types of robot systems. It helps reduce the work of the programmer. However, the applications of available systems are limited to manipulation of rigid objects. Manipulation of deformable objects is rarely considered, because it is difficult to design a method for representing states of deformable objects and operations against them. Furthermore, too many operations are possible on them. In this paper, we choose knot tying as a case study for manipulating deformable objects, because the knot theory is available and the types of operations possible in knot tying are limited. We propose a knot planning from observation (KPO) paradigm, a KPO theory, and a KPO system.
Jun Takamatsu, Takuma Morita, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
IEEE Trans. Robotics5
2005 Inverse Polarization Raytracing: Estimating Surface Shapes of Transparent Objects
abstract
We propose a novel method for estimating the surface shapes of transparent objects by analyzing the polarization state of the light. Existing methods do not fully consider the reflection, refraction, and transmission of the light occurring inside a transparent object. We employ a polarization raytracing method to compute both the path of the light and its polarization state. Our proposed iterative computation method estimates the surface shape of the transparent object by minimizing the difference between the polarization data rendered by the polarization raytracing method and the polarization data obtained from a real object.
Daisuke Miyazaki, Katsushi Ikeuchi
CVPR (2)2
2005 Reflection Components Decomposition of Textured Surfaces Using Linear Basis Functions
abstract
Most existing methods of reflection components decomposition using a single color image require color segmentation. Few methods that employ local operations are able to avoid the requirement; however, they usually suffer from color discontinuity problems. In this paper, we introduce a decomposition method using a single color image that does not require (global) color segmentation or (local) color discontinuity detection. The method principally utilizes the coefficients of the reflectance basis functions of input image and its specular-free image. Combining those coefficients enables us to find the diffuse coefficients of the specular pixels for every surface color. As a result, the decomposition becomes a well-posed problem and able to be solved in closed-form equations. Our experimental results on real complex textured images show the effectiveness of our proposed method.
Robby T. Tan, Katsushi Ikeuchi
CVPR (1)2
2005 Shape Recovery of 3D Data Obtained from a Moving Range Sensor by Using Image Sequences
abstract
For a large object, scanning from the air is one of the most efficient methods of obtaining 3D data. In the case of large cultural heritage objects, there are some difficulties in scanning with respect to safety and efficiency. To remedy these problems, we have been developing a novel 3D measurement system, the floating laser range sensor (FLRS), in which a range sensor is suspended beneath a balloon. The obtained data, however, have some distortions due to sensor-movements during the scanning process. In this paper, we propose a method to recover 3D range data obtained by a moving laser range sensor. This method is applicable not only to our FLRS, but also to a general moving range sensor. Using image sequences from a video camera mounted on the FLRS enables us to estimate the motion of the FLRS without any physical sensors such as gyros or GPS. In the first stage, the initial values of camera motion parameters are estimated by full-perspective factorization. The next stage refines camera motion parameters using the relationships between camera images and range data distortion. Finally, by using the refined parameters, the distorted range data are recovered. In addition, our method is applicable with an uncalibrated video camera and range sensor system. We applied this method to an actual scanning project, and the results showed the effectiveness of our method.
Atsuhiko Banno, Katsushi Ikeuchi
ICCV2
2005 Multiple Light Sources and Reflectance Property Estimation Based on a Mixture of Spherical Distributions
abstract
In tins paper we propose a new method for simultaneously estimating the illumination of the scene and the reflectance property of an object from a single image. We assume that the illumination consists of multiple point sources and the shape of the object is known. Unlike previous methods, we will recover not only the direction and intensity of the light sources, but also the number of light sources and the specular reflection parameter of the object. First, we represent the illumination on the surface of a unit sphere as a finite mixture of von Mises-Fisher distributions by deriving a spherical specular reflection model. Next, we estimate this mixture and the number of distributions. Finally, using this result as initial estimates, we refine the estimates using the original specular reflection model. We can use the results to render the object under novel lighting conditions
Kenji Hara, Ko Nishino, Katsushi Ikeuchi
ICCV3
2005 Consistent Surface Color for Texturing Large Objects in Outdoor Scenes
abstract
Color appearance of an object is significantly influenced by the color of the illumination. When the illumination color changes, the color appearance of the object change accordingly, causing its appearance to be inconsistent. To arrive at color constancy, we have developed a physics-based method of estimating and removing the illumination color. In this paper, we focus on the use of this method to deal with outdoor scenes, since very few physics-based methods have successfully handled outdoor color constancy. Our method is principally based on shadowed and non-shadowed regions. Previously researchers have discovered that shadowed regions are illuminated by sky light, while non-shadowed regions are illuminated by a combination of sky light and sunlight. Based on this difference of illumination, we estimate the illumination colors (both the sunlight and the sky light) and then remove them. To reliably estimate the illumination colors in outdoor scenes, we include the analysis of noise, since the presence of noise is inevitable in natural images. As a result, compared to existing methods, the proposed method is more effective and robust in handling outdoor scenes. In addition, the proposed method requires only a single input image, making it useful for many applications of computer vision
Rei Kawakami, Katsushi Ikeuchi, Robby T. Tan
ICCV2
2005 Using Extended Light Sources for Modeling Object Appearance under Varying Illumination
abstract
In this study, we demonstrate the effectiveness of using extended light sources for modeling the appearance of an object for varying illumination. Extended light sources have a radiance distribution that is similar to that of the Gaussian function and have the potential of functioning as a low-pass filter when the appearance of an object is sampled under them. This enables us to obtain a set of basis images of an object for variable illumination from input images of the object taken under those light sources without suffering aliasing caused by insufficient sampling of its appearance. Furthermore, extended light sources are useful in terms of reducing high contrast in image intensities due to specular and diffuse reflection components. This helps us observe both specular and diffuse reflection components of an object in the same image taken with a single shutter speed. We have tested our proposed approach based on extended light sources with objects of complex appearance that are generally difficult to model using image-based modeling techniques.
Imari Sato, Takahiro Okabe, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV4
2005 Motion estimation of a moving range sensor by image sequences and distorted range data
abstract
For a large scale object, scanning from the air is one of the most efficient methods of obtaining 3D data. In the case of large cultural heritage objects, there are some difficulties in scanning them with respect to safety and efficiency. To remedy these problems, we have been developing a novel 3D measurement system, the floating laser range sensor (FLRS), in which a rage sensor is suspended beneath a balloon. The obtained data, however, have some distortion due to the intra-scanning movement. In this paper, we propose a method to recover 3D range data obtained by a moving laser range sensor; this method is applicable not only to our FLRS, but also to a general moving range sensor. Using image sequences from a video camera mounted on the FLRS enables us to estimate the motion of the FLRS without any physical sensors such as gyros and GPS. At first, the initial values of camera motion parameters are estimated by perspective factorization. The next stage refines camera motion parameters using the relationships between camera images and the range data distortion. Finally, by using the refined parameter, the distorted range data are recovered. We applied this method to an actual scanning project and the results showed the effectiveness of our method.
Atsuhiko Banno, Kazuhide Hasegawa, Katsushi Ikeuchi
IROS3
2005 The climbing sensor: 3-D modeling of a narrow and vertically stalky space by using spatio-temporal range image
abstract
In this paper, we propose a novel type of 3D scanning system named 'climbing sensor'. This system has been designed for scanning narrow and vertically stalky spaces, which are hard or extremely inefficient to scan by commercial laser range scanners due to their dimensions and limitation of FOVs. The climbing sensor equips a platform with two line scanners on a lift, and they scan through the whole target while the lift moves downwards along a ladder. One scanner is for scanning the target, which scans horizontally as the lift moves vertically, and the other scanner is for localizing the platform, which scans vertically. By using spatio-temporal range image acquired from the vertical scanning, we can accurately calculate the speed of the moving platform, with which a correct 3D model can be constructed from horizontal scans. We applied this scanning system to the Bayon Temple in Cambodia as a part of our digital archiving project of cultural assets. The scanning results proved that the system gives a sufficiently accurate 3D model and the effectiveness of our proposed system and speed estimating process.
Ken Matsui, Shintaro Ono, Katsushi Ikeuchi
IROS3
2005 Task model of lower body motion for a biped humanoid robot to imitate human dances
abstract
The goal of this study is developing a biped humanoid robot that can observe a human dance performance and imitate it. To achieve this goal, we propose a task model of lower body motion, which consists of task primitives (what to do) and skill parameters (how to do it). Based on this model, a sequence of task primitives and their skill parameters are detected from human motion, and robot motion is regenerated from the detected result under constraints of a robot. This model can generate human-like lower body motion including various waist motions as well as various stepping motions of the legs. Generated motions can be performed stably on an actual robot supported by its own legs. We used improved robot hardware HRP-2, which has superior features in body weight, actuators, and DOF of the waist. By using the proposed method and HRP-2, we have realized a dance performance of Japanese folk dance by the robot, which is synchronized with a performance of a human grand master on the same stage.
Shinichiro Nakaoka, Atsushi Nakazawa, Fumio Kanehiro, Kenji Kaneko, Mitsuharu Morisawa, Katsushi Ikeuchi
IROS6
2005 Generation of humanoid robot motions with physical constraints using hierarchical B-spline
abstract
From recorded human motion, trajectory optimization of a humanoid robot with physical limits being the key constraint is presented. The optimization's objective function preserves the salient characteristics of the original motion, while constraints are used to transform the motion to the limit of the capabilities of the humanoid robot. It is shown that using constraints to limit generated trajectories to physically realizable motion ensures that limits are being met more precisely than by reducing them using a standard objective function. The use of wavelets vs. B-spline is compared, and it is shown that B-spline data representation has pronounced advantages. Furthermore, in regions of high-frequency motion not adequately modeled by an initial splining, B-spline is extensible into a hierarchy so that optimizations can be performed locally which still meet global criteria. It is shown how to use B-spline coefficients in angle-, velocity-, acceleration-, and dynamic force-constraints. To detect trajectory subsections with excessive error in need of control point adjustment and local re-optimization, not only is a traditional error detector used, but also a B-spline density detector is also presented. Generated motions were tested using our simulation program and the robot HRP-2.
Miti Ruchanurucks, Shinichiro Nakaoka, Shunsuke Kudoh, Katsushi Ikeuchi
IROS4
2005 Shading and Shadowing of Architecture in Mixed Reality
abstract
We propose a simple method to express shading and shadowing of virtual objects in mixed reality especially appropriate for static architecture models in outdoor scenes. We create the shadows of the virtual objects in a fast and efficient way using a set of pre-rendered basis images and shadowing planes. The proposed method is limited in interactivity but can operate in near real-time.
Tetsuya Kakuta, Takeshi Oishi, Katsushi Ikeuchi
ISMAR3
2005 Driving View Simulation Synthesizing Virtual Geometry and Real Images in an Experimental Mixed-Reality Traffic Space
abstract
We propose an efficient and effective image generation system for an experimental mixed-reality traffic space. Our enhanced traffic/driving simulation system represents the view through a hybrid that combines virtual geometry with real images to realize high photo-reality with little human cost. Images for datasets are captured from the real world, and the view for the simulation system is created by synthesizing image datasets - with a conventional driving simulator.
Shintaro Ono, Koichi Ogawara, Masataka Kagesawa, Hiroshi Kawasaki, Masaaki Onuki, Ken Honda, Katsushi Ikeuchi
ISMAR7
2005 Plausible image matching: determining dense and smooth mapping between images without a priori knowledge
abstract
This paper presents a method for automatic determination of dense and smooth mapping between two images without a priori knowledge of either the camera pose or the objects in the images. We designed an algorithm to find the mapping between a pair of arbitrary images, and accomplish automatic image morphing. In order to extract image features which look natural to human, we use a set of linear filters similar to those that are used in early vision. Then the derived vector fields consisting of filter responses are matched with each other through a minimization of the cost function which expresses the similarity of transformed images and mapping smoothness, in a multiresolutional hierarchy. Since the cost function in general is highly nonlinear, we avoid excessive distortion in the estimated mapping by providing a local convexity of mapping in nonlinear optimization. In this paper, a variety of experimental results are discussed for various data sets, including images of rotating objects, static objects, human faces and texture patterns, to demonstrate the performance of the proposed method.
Shuntaro Yamazaki, Katsushi Ikeuchi, Yoshihisa Shinagawa
Int. J. Pattern Recognit. Artif. Intell.2
2005 Light Source Position and Reflectance Estimation from a Single View without the Distant Illumination Assumption
abstract
Several techniques have been developed for recovering reflectance properties of real surfaces under unknown illumination. However, in most cases, those techniques assume that the light sources are located at inifinity, which cannot be applied safely to, for example, reflectance modeling of indoor environments. In this paper, we propose two types of methods to estimate the surface reflectance property of an object, as well as the position of a light source from a single view without the distant illumination assumption, thus relaxing the conditions in the previous methods. Given a real image and a 3D geometric model of an object with specular reflection as inputs, the first method estimates the light source position by fitting to the Lambertian diffuse component, while separating the specular and diffuse components by using an iterative relaxation scheme. Our second method extends that first method by using as input a specular component image, which is acquired by analyzing multiple polarization images taken from a single view, thus removing its constraints on the diffuse reflectance property. This method simultaneously recovers the reflectance properties and the light source positions by optimizing the linearity of a log-transformed Torrance-Sparrow model. By estimating the object's reflectance property and the light source position, we can freely generate synthetic images of the target object under arbitrary lighting conditions with not only source direction modification but also source-surface distance modification. Experimental results show the accuracy of our estimation framework.
Kenji Hara, Ko Nishino, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Adaptively Merging Large-Scale Range Data with Reflectance Properties
abstract
In this paper, we tackle the problem of geometric and photometric modeling of large intricately shaped objects. Typical target objects we consider are cultural heritage objects. When constructing models of such objects, we are faced with several important issues that have not been addressed in the past-issues that mainly arise due to the large amount of data that has to be handled. We propose two novel approaches to efficiently handle such large amounts of data: A highly adaptive algorithm for merging range images and an adaptive nearest-neighbor search to be used with the algorithm. We construct an integrated mesh model of the target object in adaptive resolution, taking into account the geometric and/or photometric attributes associated with the range images. We use surface curvature for the geometric attributes and (laser) reflectance values for the photometric attributes. This adaptive merging framework leads to a significant reduction in the necessary amount of computational resources. Furthermore, the resulting adaptive mesh models can be of great use for applications such as texture mapping, as we will briefly demonstrate. Additionally, we propose an additional test for the k-d tree nearest-neighbor search algorithm. Our approach successfully omits back-tracking, which is controlled adaptively depending on the distance to the nearest neighbor. Since the main consumption of computational cost lies in the nearest-neighbor search, the proposed algorithm leads to a significant speed-up of the whole merging process. In this paper, we present the theories and algorithms of our approaches with pseudo code and apply them to several real objects, including large-scale cultural assets.
Ryusuke Sagawa, Ko Nishino, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Separating Reflection Components of Textured Surfaces Using a Single Image
abstract
In inhomogeneous objects, highlights are linear combinations of diffuse and specular reflection components. A number of methods have been proposed to separate or decompose these two components. To our knowledge, all methods that use a single input image require explicit color segmentation to deal with multicolored surfaces. Unfortunately, for complex textured images, current color segmentation algorithms are still problematic to segment correctly. Consequently, a method without explicit color segmentation becomes indispensable and this paper presents such a method. The method is based solely on colors, particularly chromaticity, without requiring any geometrical information. One of the basic ideas is to iteratively compare the intensity logarithmic differentiation of an input image and its specular-free image. A specular-free image is an image that has exactly the same geometrical profile as the diffuse component of the input image and that can be generated by shifting each pixel's intensity and maximum chromaticity nonlinearly. Unlike existing methods using a single image, all processes in the proposed method are done locally, involving a maximum of only two neighboring pixels. This local operation is useful for handling textured objects with complex multicolored scenes. Evaluations by comparison with the results of polarizing filters demonstrate the effectiveness of the proposed method.
Robby T. Tan, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 A sensor fusion approach for recognizing continuous human grasping sequences using hidden Markov models
abstract
The Programming by Demonstration (PbD) technique aims at teaching a robot to accomplish a task by learning from a human demonstration. In a manipulation context, recognizing the demonstrator's hand gestures, specifically when and how objects are grasped, plays a significant role. Here, a system is presented that uses both hand shape and contact-point information obtained from a data glove and tactile sensors to recognize continuous human-grasp sequences. The sensor fusion, grasp classification, and task segmentation are made by a hidden Markov model recognizer. Twelve different grasp types from a general, task-independent taxonomy are recognized. An accuracy of up to 95% could be achieved for a multiple-user system.
Keni Bernardin, Koichi Ogawara, Katsushi Ikeuchi, Rüdiger Dillmann
IEEE Trans. Robotics3
2004 Color Alignment in Texture Mapping of Images under Point Light Source and General Lighting Condition
Hiroki Unten, Katsushi Ikeuchi
CVPR (1)2
2004 Detection of Vehicles in Panoramic Range Image
abstract
It is important to assess street-parking vehicles causing traffic problems in urban areas. However, assessments are performed manually and at high cost. Developing a detection system of those vehicles is a top priority for reducing cost and avoiding human error. We propose a detection method using a laser-range finder and a line-scan camera. In the detection method, two kinds of cluster analyses are applied to laser-range points: one is for clustering laser-range points at each scan, and the other for clustering laser-range points over several scans. Each cluster of laser-range points indicates a vehicle. As a result of verification experiments in real roads, a detection rate reached 90%.
Kiyotaka Hirahara, Katsushi Ikeuchi
ICRA2
2004 Flying Laser Range Finder and its Data Registration Algorithm
abstract
Scanning from the air is one of the most efficient methods for obtaining 3D data of large-scale objects. For this purpose, we have been developing a flying laser range finder that is suspended under a balloon. Even though the scanning speed of the finder is quite rapid, it is difficult to eliminate the influence of the swing of a balloon. As a result the scanned data have some distortion due to the intra-scanning movement. In order to compensate this intra-scanning movement, we propose a evolutional registration algorithm which not only aligns multiple range images to determine inter-scanning movement parameters, but also rectifies distortion of range image by determining intra-scanning movement parameters. In this paper, we describe our aerial scanning system especially focusing on the design of the flying laser range finder and deformation registration algorithm. To show the effectiveness of our method, we evaluate its performance using synthesized and real data.
Yuichiro Hirota, Tomohito Masuda, Ryo Kurazume, Koichi Ogawara, Kazuhide Hasegawa, Katsushi Ikeuchi
ICRA6
2004 Leg Motion Primitives for a Dancing Humanoid Robot
abstract
The goal of the study described In this work is to develop a total technology for archiving human dance motions. A key feature of this technology is a dance replay by a humanoid robot. Although human dance motions can be acquired by a motion capture system, a robot cannot exactly follow the captured data because of different body structure and physical properties between the human and the robot. In particular, leg motions are too constrained to be converted from the captured data because the legs must interact with the floor and keep dynamic balance within the mechanical constraints of current robots. To solve this problem, we have designed a symbolic description of leg motion primitives in a dance performance. Human dance actions are recognized as a sequence of primitives and the same actions of the robot can be regenerated from them. This framework is more reasonable than modifying the original motion to adapt the robot constraints. We have developed a system to generate feasible robot motions from a human performance, and realized a dance performance by the robot HRP-1S.
Shinichiro Nakaoka, Atsushi Nakazawa, Kazuhito Yokoi, Katsushi Ikeuchi
ICRA4
2004 Matching and blending human motions temporal scaleable dynamic programming
abstract
This paper presents a method for matching the frames of the human motions acquired by a motion capture system, and then creating blended (interpolated) motions according to the matching result. This matching method is basically a variation of a dynamic programming (DP) matching but we enhanced it to enable it to detect the timescale parameters. This scaleable dynamic programming (scaleable-DP) can match and evaluate the same class of motions such as walking, running, stepping and their timescale parameters. This approach is adaptable for differences in individuals, such as body sizes and timing of the stop-frames. On the blending pipeline, we first generate the keyframes according to the matching result. The keyframes are generated by considering the spatial and temporal difference of individual motions. After that, transition motions are synthesized between the keyframes. We experimented with our approach by using 15 gait motions and 5 dance motions. The results of these demonstrations show the validity of the proposed algorithm.
Atsushi Nakazawa, Shinichiro Nakaoka, Katsushi Ikeuchi
IROS3
2004 Flexible cooperation between human and robot by interpreting human intention from gaze information
abstract
This paper describes a method to realize flexible cooperation between human and robot which reflects the intention and state of human by using gaze information. This physiological information expresses the process of thinking directly, so it enables us to read the internal condition such as hesitation or search in decision making process. We propose a method to interpret the intention and condition from the latest history of gaze movement and determine an appropriate cooperative action of a robot based on it so that the task proceeds smoothly. Finally, we show experimental results by using a humanoid-type robot.
Kenji Sakita, Koichi Ogawara, Shinji Murakami, Kentaro Kawamura, Katsushi Ikeuchi
IROS5
2004 Constructing Virtual Cities by Using Panoramic Images
Katsushi Ikeuchi, Masao Sakauchi, Hiroshi Kawasaki, Imari Sato
Int. J. Comput. Vis.1
2004 Editorial: Research in Japan on Omni-Directional Sensors and Their Applications
Yasushi Yagi, Katsushi Ikeuchi
Int. J. Comput. Vis.2
2004 Illumination Normalization with Time-Dependent Intrinsic Images for Video Surveillance
abstract
Variation in illumination conditions caused by weather, time of day, etc., makes the task difficult when building video surveillance systems of real world scenes. Especially, cast shadows produce troublesome effects, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. In this paper, we handle such appearance variations by removing shadows in the image sequence. This can be considered as a preprocessing stage which leads to robust video surveillance. To achieve this, we propose a framework based on the idea of intrinsic images. Unlike previous methods of deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images instead of assuming a single reflectance image. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadowing effects. We also propose an illumination normalization scheme which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day, etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results and shows its effectiveness with object tracking results on real scene data sets.
Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Transparent Surface Modeling from a Pair of Polarization Images
Daisuke Miyazaki, Masataka Kagesawa, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Separating Reflection Components Based on Chromaticity and Noise Analysis
abstract
Many algorithms in computer vision assume diffuse only reflections and deem specular reflections to be outliers. However, in the real world, the presence of specular reflections is inevitable since there are many dielectric inhomogeneous objects which have both diffuse and specular reflections. To resolve this problem, we present a method to separate the two reflection components. The method is principally based on the distribution of specular and diffuse points in a two-dimensional maximum chromaticity-intensity space. We found that, by utilizing the space and known illumination color, the problem of reflection component separation can be simplified into the problem of identifying diffuse maximum chromaticity. To be able to identify the diffuse maximum chromaticity correctly, an analysis of the noise is required since most real images suffer from it. Unlike existing methods, the proposed method can separate the reflection components robustly for any kind of surface roughness and light direction.
Robby T. Tan, Ko Nishino, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Illumination Normalization with Time-dependent Intrinsic Images for Video Surveillance
abstract
Cast shadows produce troublesome effects for video surveillance systems, typically for object tracking from a fixed viewpoint, since it yields appearance variations of objects depending on whether they are inside or outside the shadow. To robustly eliminate these shadows from image sequences as a preprocessing stage for robust video surveillance, we propose a framework based on the idea of intrinsic images. Unlike previous methods for deriving intrinsic images, we derive time-varying reflectance images and corresponding illumination images from a sequence of images. Using obtained illumination images, we normalize the input image sequence in terms of incident lighting distribution to eliminate shadow effects. We also propose an illumination normalization scheme, which can potentially run in real time, utilizing the illumination eigenspace, which captures the illumination variation due to weather, time of day etc., and a shadow interpolation method based on shadow hulls. This paper describes the theory of the framework with simulation results, and shows its effectiveness with object tracking results on real scene data sets for traffic monitoring.
Yasuyuki Matsushita, Ko Nishino, Katsushi Ikeuchi, Masao Sakauchi
CVPR (1)3
2003 Illumination Chromaticity Estimation using Inverse-Intensity Chromaticity Space
abstract
Existing color constancy methods cannot handle both uniform colored surfaces and highly textured surfaces in a single integrated framework. Statistics-based methods require many surface colors, and become error prone when there are only few surface colors. In contrast, dichromatic-based methods can successfully handle uniformly colored surfaces, but cannot be applied to highly textured surfaces since they require precise color segmentation. In this paper, we present a single integrated method to estimate illumination chromaticity from single/multi-colored surfaces. Unlike the existing dichromatic-based methods, the proposed method requires only rough highlight regions, without segmenting the colors inside them. We show that, by analyzing highlights, a direct correlation between illumination chromaticity and image chromaticity can be obtained. This correlation is clearly described in "inverse-intensity chromaticity space", a new two-dimensional space we introduce. In addition, by utilizing the Hough transform and histogram analysis in this space, illumination chromaticity can be estimated robustly, even for a highly textured surface. Experimental results on real images show the effectiveness of the method.
Robby T. Tan, Ko Nishino, Katsushi Ikeuchi
CVPR (1)3
2003 Determining Reflectance and Light Position from a Single Image Without Distant Illumination Assumption
abstract
Several techniques have been developed for recovering reflectance properties of real surfaces under unknown illumination conditions. However, in most cases, those techniques assume that the light sources are located at infinity, which cannot be applied to, for example, photometric modelling of indoor environments. We propose two methods to estimate the surface reflectance property of an object, as well as the position of a light source from a single image without the distant illumination assumption. Given a color image of an object with specular reflection as an input, the first method estimates the light source position by fitting to the Lambertian diffuse component, while separating the specular and diffuse components by using an iterative relaxation scheme. Moreover, we extend the above method by using a single specular image as an input, thus removing its constraints on the diffuse reflectance property and the number of light sources. This method simultaneously recovers the reflectance properties and the light source positions by optimizing the linearity of a log-transformed Torrance-Sparrow model. By estimating the object's reflectance property and the light source position, we can freely generate synthetic images of the target object under arbitrary source directions and source-surface distances.
Kenji Hara, Ko Nishino, Katsushi Ikeuchi
ICCV3
2003 Polarization-based Transparent Surface Modeling from Two Views
abstract
In this paper, we propose a novel method to recover the surface shape of transparent objects. The degree of polarization of the light reflected from the object surface depends on the reflection angle which, in turn, depends on the object's surface normal; thus, by measuring the degree of polarization, we are able to calculate the surface normal of the object. However, degree of polarization and surface normal does not correspond one-to-one, making us to analyze two polarization images taken from two different view in order to solve the ambiguity. A parabolic curve will be a strong clue to correspond a point in one image to a point in the other image, where both points represent the same point on object surface. By comparing the degree of polarization at such corresponding points, the true surface normal can be determined.
Daisuke Miyazaki, Masataka Kagesawa, Katsushi Ikeuchi
ICCV3
2003 Polarization-based Inverse Rendering from a Single View
abstract
This paper presents a method to estimate geometrical, photometrical, and environmental information of a single-viewed object in one integrated framework under fixed viewing position and fixed illumination direction. These three types of information are important to render a photorealistic image of a real object. Photometrical information represents the texture and the surface roughness of an object, while geometrical and environmental information represent the 3D shape of an object and the illumination distribution, respectively. The proposed method estimates the 3D shape by computing the surface normal from polarization data, calculates the texture of the object from the diffuse only reflection component, determines the illumination directions from the position of the brightest intensity in the specular reflection component, and finally computes the surface roughness of the object by using the estimated illumination distribution.
Daisuke Miyazaki, Robby T. Tan, Kenji Hara, Katsushi Ikeuchi
ICCV4
2003 Appearance Sampling for Obtaining A Set of Basis Images for Variable Illumination
abstract
Previous studies have demonstrated that the appearance of an object under varying illumination conditions can be represented by a low-dimensional linear subspace. A set of basis images spanning such a linear subspace can be obtained by applying the principal component analysis (PCA) for a large number of images taken under different lighting conditions. While the approaches based on PCA have been used successfully for object recognition under varying illumination conditions, little is known about how many images would be required in order to obtain the basis images correctly. In this study, we present a novel method for analytically obtaining a set of basis images of an object for arbitrary illumination from input images of the object taken under a point light source. The main contribution of our work is that we show that a set of lighting directions can be determined for sampling images of an object depending on the spectrum of the object's BRDF in the angular frequency domain such that a set of harmonic images can be obtained analytically based on the sampling theorem on spherical harmonics. In addition, unlike the previously proposed techniques based on spherical harmonics, our method does not require the 3D shape and reflectance properties of an object used for rendering harmonics images of the object synthetically.
Imari Sato, Takahiro Okabe, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV4
2003 Separating Reflection Components of Textured Surfaces using a Single Image
abstract
The presence of highlights, which in dielectric inhomogeneous objects are linear combination of specular and diffuse reflection components, is inevitable. A number of methods have been developed to separate these reflection components. To our knowledge, all methods that use a single input image require explicit color segmentation to deal with multicolored surfaces. Unfortunately, for complex textured images, current color segmentation algorithms are still problematic to segment correctly. Consequently, a method without explicit color segmentation becomes indispensable, and this paper presents such a method. The method is based solely on colors, particularly chromaticity, without requiring any geometrical parameter information. One of the basic ideas is to compare the intensity logarithmic differentiation of specular-free images and input images iteratively. The specular-free image is a pseudo-code of diffuse components that can be generated by shifting a pixel's intensity and chromaticity nonlinearly while retaining its hue. All processes in the method are done locally, involving a maximum of only two pixels. The experimental results on natural images show that the proposed method is accurate and robust under known scene illumination chromaticity. Unlike the existing methods that use a single image, our method is effective for textured objects with complex multicolored scenes.
Robby T. Tan, Katsushi Ikeuchi
ICCV2
2003 Knot planning from observation
abstract
Learning from Observation (LFO) has been widely applied in various types of robot system. It helps reduce the work of the programmer. But the available systems have application limited to rigid objects. Deformable objects are not considered because: 1) it is difficult to describe their state, and 2) too many operations are possible on them. In this paper, we choose the knot tying as case study for operating on nonrigid bodies, because a "knot theory" is available and the type of operations is limited. We describe the Knot Planning from Observation (KPO) paradigm, KPO theory and KPO system.
Takuma Morita, Jun Takamatsu, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
ICRA5
2003 Generating whole body motions for a biped humanoid robot from captured human dances
abstract
The goal of this study is a system for a robot to imitate human dances. This paper describes the process to generate whole body motions which can be performed by an actual biped humanoid robot. Human dance motions are acquired through a motion capturing system. We then extract symbolic representation which is made up of primitive motions: essential postures in arm motions and step primitives in leg motions. A joint angle sequence of the robot is generated according to these primitive motions. Then joint angles are modified to satisfy mechanical constraints of the robot. For balance control, the waist trajectory is moved to acquire dynamics consistency based on desired ZMP. The generated motion is tested on OpenHRP dynamics simulator. In our test, the Japanese folk dance, 'Jongara-bushi', was successfully performed by HRP-1S.
Shinichiro Nakaoka, Atsushi Nakazawa, Kazuhito Yokoi, Hirohisa Hirukawa, Katsushi Ikeuchi
ICRA5
2003 Synthesize stylistic human motion from examples
abstract
The human body motion synthesis is highly necessary for humanoid robots' motion planning and computer animations. In this paper, new method for generating human-like natural motions based on the motion database acquired by motion capture systems is described. On the analysis step, the acquired motions are divided into some motion segments, and then the characteristic poses and motions are archived as 'motion styles'. The motion style is a kind of the human skill, and it's unique to the motions' scenario, such as the different kinds of dances. On the synthesis step, users direct the key poses of human figures. The system generates the characteristic motions according to the user's directions and motion style database. The experiment result shows that this method can synthesize the realistic 'stylized' motions with this framework.
Atsushi Nakazawa, Shinichiro Nakaoka, Katsushi Ikeuchi
ICRA3
2003 Estimation of essential interactions from multiple demonstrations
abstract
To learn a new everyday task under the "Learning from Observation" framework, the system needs to detect which parts of the demonstration are essential to complete the task without task-dependent knowledge. In the previous research, we proposed a technique to estimate essential interactions in a task by integrating multiple demonstrations which represent virtually the same task. Although, the technique could automatically segment the essential interactions and determine the number of the interactions, the segmentation algorithm depends on some heuristics and only stationary interactions could be obtained. In this paper, a novel technique is proposed, which overcomes this limitation and can estimate almost any types of interactions. In this approach, a demonstrator needs to give a explicit signal once during each essential interaction as a hint on the occurrence of the essential interaction. From visual information and these signals, the system automatically analyzes the essential parts of the task and their periods, and also detects which environmental objects are interacted with the manipulated object. These information is hard to be obtained from a single demonstration, because of the ambiguity in interpreting the interaction especially in cluttered environment. The proposed method is evaluated in a simulation and also in a real world by using a humanoid robot.
Koichi Ogawara, Jun Takamatsu, Hiroshi Kimura, Katsushi Ikeuchi
ICRA4
2003 Calculating possible local displacement of curve objects using improved screw theory
abstract
Various methods to recognize assembly tasks using possible local displacement of objects have been proposed. To calculate this displacement, the screw theory is employed. It is equivalent to the first order Taylor expansion of the displacement. However, such methods can treat polyhedral objects only. Because the screw theory cannot treat curvature information of objects. In this paper, we propose a method to calculate possible local displacement of curve objects using improved screw theory, which is equivalent to the second order Taylor expansion of the displacement, and verify the validity of the proposed method.
Jun Takamatsu, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
ICRA4
2003 Grasp recognition using a 3D articulated model and infrared images
abstract
A technique to recognize the shape of a grasping hand during manipulation tasks is proposed; which utilizes 3D articulated hand model and a reconstructed 3D volume from infrared cameras. Vision-based recognition of a grasping hand is a tough problem, because a hand may be partially occluded by a grasped object and the ratio of occlusion changes along the progress of the task. To recognize the shape in a single time frame, a robust recognition method of an articulated object is proposed. In this method, 3D volumetric representation of a hand is reconstructed from multiple silhouette images and 3D articulated object model is fitted to be reconstructed data to estimate the pose and the joint angles. To deal with large occlusion, a technique to simultaneously estimate time series reconstructed volumes with the above method is proposed, which can automatically suppress the effect form badly reconstructed volumes. The proposed techniques are verified in simulation as well as in a real world.
Koichi Ogawara, Jun Takamatsu, Kentaro Hashimoto, Katsushi Ikeuchi
IROS4
2003 The Great Buddha Project: Modeling Cultural Heritage for VR Systems through Observation
Katsushi Ikeuchi, Atsushi Nakazawa, Kazuhide Hasegawa, Takeshi Oishi
ISMAR1
2003 Sensing requirements for robotic assembly from an analysis of critical contact transitions
abstract
This paper presents a method for determining sensing requirements for robotic assembly from an analysis of critical contact-state transitions produced among mating parts during the execution of nominal assembly plans. The goal is to support the reduction of real-life uncertainty through the recognition of assembly tasks that require force and visual feedback operations. The assembly tasks are decomposed into assembly skill primitives based on transitions described on a taxonomy of contact relations. Force feedback operations are described as a set of force compliance skills which are systematically associated to the assembly skill primitives. To determine the visual feedback operations and the type of visual information needed, a backward propagation process of geometrical constraints is used. This process defines new visual feedback requirements for the tasks from the discovery of direct, and indirect, insertion and contact dependencies among the mating parts.
Santiago E. Conant-Pablos, Horacio Martinez-Alfaro, Katsushi Ikeuchi
SMC3
2003 Hardware-accelerated visualization of volume-sampled distance fields
abstract
We present a method of visualizing volume-sampled distance fields, taking advantage of hardware-acceleration of modern graphics hardware. Although conventional distance fields can represent only 2-manifold surfaces in a stable way, we have developed a technique of representing both manifold and non-manifold surfaces in a volume created by sampling a segmented distance field. The volume-sampled distance fields can be visualized effectively with pre-integrated volume rendering by embedding the interpolating function as a lookup table called vertex generation diagram into graphics hardware. We developed a simple system for smoothly blending two surface models in the distance function domain, and confirmed that our proposed representation of surface models is applicable to existing methods of geometric processing using implicit representations. We also confirmed that proposed methods of deformation and visualization can be executed at sufficient quality and speed using commodity graphics hardware on a standard PC.
Shuntaro Yamazaki, Kiwamu Kase, Katsushi Ikeuchi
Shape Modeling International3
2003 Shape difference visualization for ancient bronze mirrors through 3D range images
abstract
Abstract Japanese archaeologists have paid special attention to ancient Chinese bronze mirrors because the mirrors may provide a key for the exact location of Yamatai State, which is one of the major archaeological controversies. Currently, archaeologists visually analyse ancient Chinese bronze mirrors for their shape difference. The practice requires a huge amount of time and effort. In this paper, we propose an automatic method for detecting the shape difference between a pair of ancient mirrors. The 3D data of the mirrors are obtained using a laser range scanner. Our algorithm then aligns them into the same coordinate and visualizes their shape differences. Our proposed algorithm provides fast and non‐damaging analysis for shape difference. Further analysis can be evaluated on our data instead of the actual mirror, so it can be performed by more than one group of archaeologists. Copyright © 2003 John Wiley & Sons, Ltd.
Tomohito Masuda, Setsuo Imazu, Supatana Auethavekiat, Tsuyoshi Furuya, Kunihiko Kawakami, Katsushi Ikeuchi
Comput. Animat. Virtual Worlds6
2003 Illumination from Shadows
abstract
In this paper, we introduce a method for recovering an illumination distribution of a scene from image brightness inside shadows cast by an object of known shape in the scene. In a natural illumination condition, a scene includes both direct and indirect illumination distributed in a complex way, and it is often difficult to recover an illumination distribution from image brightness observed on an object surface. The main reason for this difficulty is that there is usually not adequate variation in the image brightness observed on the object surface to reflect the subtle characteristics of the entire illumination. In this study, we demonstrate the effectiveness of using occluding information of incoming light in estimating an illumination distribution of a scene. Shadows in a real scene are caused by the occlusion of incoming light and, thus, analyzing the relationships between the image brightness and the occlusions of incoming light enables us to reliably estimate an illumination distribution of a scene even in a complex illumination environment. This study further concerns the following two issues that need to be addressed. First, the method combines the illumination analysis with an estimation of the reflectance properties of a shadow surface. This makes the method applicable to the case where reflectance properties of a surface are not known a priori and enlarges the variety of images applicable to the method. Second, we introduce an adaptive sampling framework for efficient estimation of illumination distribution. Using this framework, we are able to avoid a unnecessarily dense sampling of the illumination and can estimate the entire illumination distribution more efficiently with a smaller number of sampling directions of the illumination distribution. To demonstrate the effectiveness of the proposed method, we have successfully tested the proposed method by using sets of real images taken in natural illumination conditions with different surface materials of shadow regions.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 Non-Manifold Implicit Surfaces Based on Discontinuous Implicitization and Polygonization
abstract
Implicit surfaces in 3D geometric modeling are limited to two manifolds because the corresponding implicit fields are usually defined by real-valued functions which bisect space into interior and exterior. We present a novel method of modeling non-manifold surfaces by implicit representation. Our method allows discontinuity of the field function and assesses the special meaning of the locus where the function is not differentiable. The enhancement can yield a non-manifold surface with such features as holes and boundaries. The discontinuous field function also enables multiple classification of the field, which makes it possible to represent branches and intersections of the implicit surfaces. The implicit field is polygonized by the algorithm based on the marching cubes algorithm, which is extended to treat discontinuous fields correctly. We also describe an efficient implementation of converting a surface model into a set of discrete samples of field function, and finally present the result of the non-manifold surfaces reproduced by our method.
Shuntaro Yamazaki, Kiwamu Kase, Katsushi Ikeuchi
GMP3
2002 Generation of a Task Model by Integrating Multiple Observations of Human Demonstrations
abstract
This paper describes a new approach on how to teach a robot everyday manipulation tasks under the "learning from observation" framework. Most of the approaches so far assume that a demonstration can be well understood from a single demonstration. However, a single demonstration contains ambiguity, in that interactions which are essential to complete a task cannot be discerned without prior task dependent knowledge, which should be obtained from observation. To address these issues, we propose a technique to integrate multiple observations of demonstrations. The demonstrations differ, but are virtually the same task. The shared interactions among all the demonstrations are considered to be essential and we form a task model from their symbolic representations. Then the relative trajectories corresponding to each essential interaction are generalized by calculating their mean and variance and are also stored in the task model, which is used to reproduce a skilled behavior. We examine this approach by using a human-form robot, which successfully imitates human demonstrations of everyday tasks.
Koichi Ogawara, Jun Takamatsu, Hiroshi Kimura, Katsushi Ikeuchi
ICRA4
2002 The dynamic postural adjustment with the quadratic programming method
abstract
The postural balance system is one of the most fundamental functions for humanoid robot control. In this paper, we propose a new feedback balance control system for the human body. This system can manipulate large perturbations. It finds the optimal motion for maintaining balance in the 3D space without receiving any feed-forward input beforehand. Two different strategies are adopted for the optimization: the quadratic programming method and the PD control. Simulation results are compared with real human motion; many common features such as rotating arms are observed.
Shunsuke Kudoh, Taku Komura, Katsushi Ikeuchi
IROS3
2002 Imitating human dance motions through motion structure analysis
abstract
This paper presents a method for importing human dance motion into humanoid robots through visual observation. The human motion data is acquired from a motion capture system consisting of 8 cameras and 8 PC clusters. Then the whole motion sequence is divided into motion elements and clustered into groups according to the correlation of end-effector trajectories. We call these segments 'motion primitives'. New dance motions are generated by concatenating these motion primitives. We are also trying to make a humanoid dance these original or generated motions using inverse-kinematics and dynamic balancing techniques.
Atsushi Nakazawa, Shinichiro Nakaoka, Katsushi Ikeuchi, Kazuhito Yokoi
IROS3
2002 Modeling manipulation interactions by hidden Markov models
abstract
This paper describes a new approach on how to teach everyday manipulation tasks to a robot under the "Learning from Observation" framework. In our previous work, to acquire low-level action primitives of a task automatically, we proposed a technique to estimate essential interactions to complete a task by integrating multiple observations of similar demonstrations. But after many demonstrations are performed, there may be interactions which are the same in nature. These identical interactions should be grouped so that each action primitive becomes unique. For this purpose, a Hidden Markov Model based clustering algorithm is presented which automatically determines the number of independent interactions. We also show that the obtained interactions can be used as discriminators of human behavior. Finally, simulation and experimental results in which a real humanoid robot learns and recognizes essential actions by observing demonstrations are presented.
Koichi Ogawara, Jun Takamatsu, Hiroshi Kimura, Katsushi Ikeuchi
IROS4
2002 Iterative refinement of range images with anisotropic error distribution
abstract
We propose a method which refines the range measurement of range finders by computing correspondences of vertices of multiple range images acquired from various viewpoints. Our method assumes that a range image acquired by a laser rangefinder has anisotropic error distribution which is parallel to the ray direction. Thus, we find the corresponding points of range images along with the ray direction. We iteratively converge range images to minimize the distance of corresponding points. We demonstrate the effectiveness of our method by presenting the experimental results of artificial and real range data. Also, we show that our method refines a 3D shape more accurately as opposed to that achieved by using the Gaussian filter.
Ryusuke Sagawa, Takeshi Oishi, Atsushi Nakazawa, Ryo Kurazume, Katsushi Ikeuchi
IROS5
2002 Task analysis based on observing hands and objects by vision
abstract
In order to transmit, share and store human knowledge, it is important for a robot to be able to acquire task models and skills by observing human actions. Since vision plays an important role for observation, we propose a technique for measuring the position and posture of objects and hands in 3-dimensional space at high speed and with high precision by vision. Next, we show a framework for the analysis and description of a task by using an object functions as elements.
Yoshihiro Sato, Keni Bernardin, Hiroshi Kimura, Katsushi Ikeuchi
IROS4
2002 Calculating optimal trajectories from contact transitions
abstract
The learning-from-demonstration method is considered for a novel robot-programming style. It consists of two parts: 1) to recognize human performance from observation as sequential motion primitives; and 2) to execute the same performance. We (2000) have proposed a method to recognize assembly tasks. However, the execution requires the ability to convert motion primitives to collision free paths. In this paper, we describe a method to calculate collision free paths. Many researchers have proposed to calculate collision free paths using analytical methods, potential fields or probabilistic methods. Potential and probabilistic methods are very powerful tools on a computer, but their solutions are not optimal. We propose a method to calculate optimal collision free paths analytically.
Jun Takamatsu, Hiroshi Kimura, Katsushi Ikeuchi
IROS3
2002 Improved screw theory using second order terms
abstract
The local displacement of an object is very useful for deciding grasp stability, generating trajectories, recognizing assembly tasks, and so on. To calculate this displacement, the screw theory is employed. It is equivalent to the first order Taylor expansion of the displacement. The screw theory is very convenient, because the displacement is formulated as simultaneous linear inequalities, and a powerful tool to calculate such inequalities, the theory of the polyhedral convex cones, has already been established. However, truncation errors introduced by first order approximations sometimes cause mistaken results. In this paper, we improve the screw theory by using 2nd order terms and verify the validity of the result.
Jun Takamatsu, Hiroshi Kimura, Katsushi Ikeuchi
IROS3
2002 Correcting observation errors for assembly task recognition
abstract
The completion of robot programs requires long development time and much effort. To shorten the programming time and to minimize the effort, we have been developing a system which we refer to as the "assembly-plan-from-observation (APO) system." This system requires assembly task recognition from observing human performance. Observation data by a robot's vision system is usually error contaminated, and, thus, we cannot use those data directly. This paper proposes two methods to clean up those errors by using contact relations and their transitions. The first one corrects the observed configuration from contact relations observed. The second one identifies wrongly determined contact relations from an analysis of configuration space (C-space). We have implemented both methods on our test bed and have verified their effectiveness.
Jun Takamatsu, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
IROS4
2002 Interactive Visualization of Non-Manifold Implicit Surfaces Using Pre-Integrated Volume Rendering
abstract
We present an interactive method of visualizing both manifold and non-manifold implicit surfaces. The implicit surfaces are directly visualized at interactive frame rates independent of surface complexity by the hardware-accelerated volume rendering method. Although conventional implicit surfaces can represent only two-manifold surfaces, we have developed a technique of representing non-manifold surfaces as implicit surfaces by applying segmented distance field. Implicit surfaces represented by this definition can be rendered by pre-integrated volume rendering using vertex generation diagrams as pre-integration table. We also developed a system for visualizing the implicit surfaces and have confirmed that it can render surfaces at sufficient quality and speed.
Shuntaro Yamazaki, Kiwamu Kase, Katsushi Ikeuchi
PG3
2002 Guest editorial
Katsushi Ikeuchi, Masataka Kagesawa, Shunsuke Kamijo
IEEE Trans. Intell. Transp. Syst.1
2001 Light Field Rendering for Large-Scale Scenes
abstract
In this paper, we present an efficient method to synthesize large-scale scenes, such as broad city landscapes. To date, model based approaches have mainly been adopted for this purpose, and some fairly convincing polygon cities have been successfully generated. However, the shapes of real world objects are usually very complicated and it is infeasible to model an entire city realistically. On the other hand, image based approaches have been attempted only recently. Image based methods are effective for realistic rendering, but their huge data sets and restrictions on interactivity pose serious problems for an actual application. Thus, we propose a hybrid method, which uses simple shapes such as planes to model the city, and applies image based techniques to add realism. It can be performed automatically through a simple image capturing process. Further, we also analyze the relationship between error and number of needed images to reduce the data size.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR (2)2
2001 Robust and Adaptive Integration of Multiple Range Images with Photometric Attributes
abstract
Integration of multiple range images is important to make use of 3D data acquired from stereo systems, laser range finders, etc. We propose a new range image integration method based on volumetric representation. Unlike other volume-based integration methods, we adaptively subdivide voxels depending on the curvature of the surface to be reconstructed, providing efficient representation of the underlying geometry and efficient use of computational resources. In our range image merging framework, additional attributes, e.g., color, laser reflectance power, etc., can be taken into account as well as 3D geometric information. This ability allows us to generate 3D models preserving sharp edges around texture boundaries, thereby providing a good basis for efficient rendering and texture mapping. The overall framework is designed to be robust against noise, taking consensus carefully in both geometry and color, which could be suitable for 3D model reconstruction from noisy stereo images. In this paper, we describe the system, and present several results of applying our framework to real data. We also present some other future applications based on our framework.
Ryusuke Sagawa, Ko Nishino, Katsushi Ikeuchi
CVPR (2)3
2001 Stability Issues in Recovering Illumination Distribution from Brightness in Shadows
abstract
The paper describes a robust method for estimating, in a reliable manner, the illumination distribution of a real scene from shadows in a given image. In general, shadows in a scene are caused by the occlusion of incoming light; image brightness inside shadows have great potential for providing distinct clues to the illumination distribution of the scene. Taking advantage of this fact, we recently proposed to estimate the illumination distribution of a real scene from a single image of the scene. The proposed method has been applied successfully to real images with complex illumination distributions. Nevertheless, it was found that sometimes the method failed to provide a correct estimate of illumination distribution. Those failures stem from the fact that the method does not take into account several factors regarding the stability of illumination estimation. The study analyzes how much information is obtainable from a given image about the illumination distribution of the scene. In particular, we carefully examine the source of instability of using shadows obtained from a single image for the estimation in several aspects: blocked view of shadows by the object; limited sampling resolution for image brightness inside shadows; and the appropriate light model to approximate the illumination distribution of the scene. Based on this analysis, we propose a method that guarantees to reliably estimate the illumination distribution of a scene, regardless of the type of input image.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR (2)3
2001 Determining Reflectance Parameters and Illumination Distribution from a Sparse Set of Images for View-dependent Image Synthesis
Ko Nishino, Zhengyou Zhang, Katsushi Ikeuchi
ICCV3
2001 Image-based rendering for mixed reality
abstract
In this paper, we propose an image-based approach to synthesize a novel view image for mixed reality (MR) systems. Theoretically, the image-based method is good for synthesizing realistic images, but it is difficult to achieve interactive handling of the object. As a solution, we propose a new method based on the "surface light field rendering" technique. With this method, we can synthesize the objects with arbitrary deformation and illumination changes. To demonstrate the efficiency of this method, we describe successful experiments that we performed using objects with non-rigid effects (e.g. velvet and Tatami carpet) which are difficult to render correctly by the use of general model-based rendering techniques.
Hiroshi Kawasaki, Hiroyuki Aritaki, Katsushi Ikeuchi, Masao Sakauchi
ICIP (3)3
2001 Automated System of Acquiring and Visualizing Tra.c Event Statistics from Tra.c Images
abstract
One of the most important application on Intelligent Transporting System (ITS) is to analyze various traffic activities and construct traffic monitoring system. However, such analyses in previous works have been done by manual inspection to huge amount of traffic images. The major reason why automated analyses of traffic images have been failed is that there does not exist any robust tracking algorithms against such crowded situations at intersections. In order to resolve such a problem, we have developed the tracking algorithm based on Spatio-Temporal Markov Random Field model which is robust against occlusion and clutter problems. Since the algorithm is able to tracking each individual vehicle even in cases of heavy occlusion situations in crowded traffic, this enables to automatically acquire detailed information from traffic images. Therefore utilizing this tracking algorithm, we then constructed an automated system which is able to acquire traffic event statistics such as vehicle counts distinguishing travel directions, velocities, frequent paths and so on. Besides, this system integrates such various statistical information in order to display correlations among different statistics clearly.
Tsunetoshi Nishida, Shunsuke Kamijo, Katsushi Ikeuchi
ICME3
2001 Acquiring Hand-action Models by Attention Point Analysis
abstract
This paper describes our current research on learning task level representations by a robot through observation of human demonstrations. We focus on human hand actions and represent such hand actions in symbolic task models. We propose a framework of such models by efficiently integrating multiple observations based on attention points; we then evaluate the model by using a human-form robot. We propose a two-step observation mechanism. At the first step, the system roughly observes the entire sequence of the human demonstration, builds a rough task model and extracts attention points (APs). The attention points indicate the time and position in the observation sequence that requires further detailed analysis. At the second step, the system closely examines the sequence around the APs and the obtained attribute values for the task model, such as what to grasp, which hand to be used, or what is the precise trajectory of the manipulated object. We implemented this system on a human form robot and demonstrated its effectiveness.
Koichi Ogawara, Tomikazu Tanuki, Hiroshi Kimura, Katsushi Ikeuchi
ICRA4
2001 Refining hand-action models through repeated observations of human and robot behavior by combined template matching
abstract
Describes research on learning task level representations by a robot through observation of human demonstrations. We focus on human hand actions and develop a construction method of a human task model which integrates multiple observations to solve ambiguity based on attention points (AP). So far, this analysis constructs a symbolic task model efficiently in a coarse-to-fine way through two steps. However, to represent delicate motion appearing in a task, the system must incorporate the information about precise motion of the manipulated objects into an abstract task model. We propose a method to identify the manipulated object through repeated observations of both human and robot behavior. To this end, we present a method which combines 2D and 3D template matching techniques to localize an object in 3D space generated from a depth and an intensity image. We apply this technique to recognition of human and robot behavior by obtaining the precise trajectory of the manipulated objects. We also present the experimental results achieved through the use of a human-form robot equipped with a 9-eye stereo vision system.
Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
IROS3
2001 Parallel processing of range data merging
abstract
This paper describes a volumetric view-merging algorithm that generates a consensus surface of an object from its range images. Our original method merges a set of range images into a volumetric implicit-surface representation, which is converted to a surface mesh by using a variant of the marching-cubes algorithm. We propose a method that increases the computation and memory efficiency for computing signed distances and the method of parallel computing on a PC cluster Since our method permits a reduction in the data amount allocated in memory, the closest point is searched efficiently; this allows us to increase the number of parallel traversals and to reduce the computation time. In this paper, we describe the following two algorithms which are complementary in terms of the efficiency of CPU and memory usage: distributed allocation of range data and parallel traversal of partial octrees. By adjusting them according to the system specifications, we can build the model efficiently by a PC cluster We have implemented this system and evaluated its performance.
Ryusuke Sagawa, Ko Nishino, Mark D. Wheeler, Katsushi Ikeuchi
IROS4
2001 Eigen-Texture Method: Appearance Compression and Synthesis Based on a 3D Model
abstract
Image-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. However, both methods still have several drawbacks when we attempt to apply them to mixed reality where we integrate virtual images with real background images. To overcome these difficulties, we propose a new method, which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface generated from a sequence of range images. The Eigen-Texture method is an example of a view-dependent texturing approach which combines the advantages of image-based and model-based approaches. No reflectance analysis of the object surface is needed, while an accurate 3D geometric model facilitates integration with other scenes. The paper describes the method and reports on its implementation.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Recognizing vehicles in infrared images using IMAP parallel vision board
abstract
Describes a method for vehicle recognition, in particular for recognizing make and model. Our system takes into account the fact that vehicles of the same make and model number come in different colors; it employs infrared (IR) images, thereby eliminating color differences. The use of IR images also enables us to use the same algorithm both day and night. This ability is particularly important because the algorithm must be able to locate many feature points, especially at night. Our algorithm is based on a configuration of local features. For the algorithm, our system first makes a compressed database of local features of a target vehicle from training images given in advance; the system then matches a set of local features in the input image with those in the training images for recognition. This method has the following three advantages: (1) it can detect even if part of the target vehicle is occluded; (2) it can detect even if the target vehicle is translated due to running out of the lanes; and (3) it does not require us to segment a vehicle part from input images. We have two implementations of the algorithm. One is referred to as the eigenwindow method, while the other is called the vector-quantization method. The former method is good at recognition, but is not very fast. The latter method is not very good at recognition but it is suitable for an IMAP parallel image-processing board; hence, it can be fast. In both implementations, the above-mentioned advantages have been confirmed by performing outdoor experiments.
Masataka Kagesawa, Shinichi Ueno, Katsushi Ikeuchi, Hiroshi Kashiwagi
IEEE Trans. Intell. Transp. Syst.3
2000 Spatio-Temporal Analysis of Omni Image
abstract
This paper describes an efficient method to obtain 3D information by using spatio-temporal analysis of omni images for outdoor navigation and map-making in the intelligent transportation system (ITS) application. Two types of omni-directional cameras are employed to make a spatio-temporal volume, which is a sequence of omni images stacked in the spatio-temporal space. For the spatio-temporal analysis of an omni image, we define several different cross sections in such spatio-temporal volumes, and examine characteristics of the traces of image features on the cross sections. We determine that the vertical straight lines in the real world are preserved as straight lines on these cross sections and that the degree of this slope represents the quotient of the velocity of the camera motion and the depth of the object. To acquire 3D information using these characteristics, we propose a hybrid method of the epipolar-plane image (EPI) analysis and the model-based analysis. To demonstrate the effectiveness of this method, we present some experimental results and the ITS applications using an omni-directional video camera to obtain images in outdoor environments.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR2
2000 Arbitrary View Position and Direction Rendering for Large-Scale Scenes
abstract
This paper presents a new method for rendering views, especially those of large-scale scenes, such as broad city landscapes. The main contribution of our method is that we are able to easily render any view from an arbitrary point to an arbitrary direction on the ground in a virtual environment. Our method belongs to the family of work that employs plenoptic functions; however, unlike other works of this type, this particular method allows us to render a novel view from almost any point on the plane at which images are taken. Previous methods, on the other hand, have some restraints concerning their re-constructable area. Thus, when synthesizing a large-scale virtual environment such as a city, our method has a great advantage. One of the applications of our method is a driving simulator in the ITS domain. We can generate any view on any lane on the road from images taken by running along just one lane. Our method, using an omni-directional camera or a measuring device of a similar type, first captures panoramic images by running along a straight line, recording the capturing position of each image. When rendering, the method divides the stored panoramic images into vertical slits, selects some suitable ones based on our theory, and reassembles them for generating an image. The method can make a virtual city with walk-through capabilities. In that virtual city, people can move and look rather freely. In this paper, we describe the basic theory of a new plenoptic function, analyze the applicable areas of the theory and the characteristics of generated images, and demonstrate a complete working system using both indoor and outdoor scenes.
Takuji Takahashi, Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR3
2000 Occlusion Robust Tracking Utilizing Spatio-Temporal Markov Random Field Model
abstract
It is very important to achieve reliable vehicle tracking in ITS application such as accident detection. The most difficult problem associated with vehicle tracking is the occlusion effect among vehicles. In order to resolve this problem, we applied the dedicated algorithm which we defined as spatio-temporal Markov random field model to traffic images at an intersection. The spatio-temporal MRF considers texture correlations between consecutive images as well as the correlation among neighbors within a image. As a result, we were able to track vehicles at the intersection robustly against occlusions. Vehicles appear in various kinds of shapes and they move in random manners at the intersection. Although occlusions occur in such complicated manners, the algorithm given was able to segment and track such occluded vehicles at a high success rate of 93-96%. The algorithm requires only gray scale images and does not assume any physical models of vehicles.
Shunsuke Kamijo, Yasuyuki Matsushita, Katsushi Ikeuchi, Masao Sakauchi
ICPR3
2000 EPI Analysis of Omni-Camera Image
abstract
The paper describes an efficient method to obtain 3D information from omni-camera images using epipolar-plane image (EPI) analysis. Two types of omni cameras are employed to make a spatio-temporal volume, which is a sequence of omni images stacked in the spatio-temporal space. For the EPI analysis of omni image, we examine different types of cross sections in such spatio-temporal volumes. To conduct the EPI analysis realisticly, we must find out the cross section on which the vertical straight lines in the real world are preserved as straight lines. We define such a cross section as omni EPI. To acquire 3D information using the characteristics of the omni EPI, we propose model based EPI analysis. To demonstrate the effectiveness of this method, we present experimental results using an omni video of outdoor environments.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
ICPR2
2000 Localization of Objects in Electric Distribution Systems by Using Segmentation and 3D Template Matching with M-Estimators
abstract
Semi-automatic mobile-robots for maintaining works in electric distribution systems have been developed by Kyushu Electric Power Company in Japan. This robot can greatly reduce the demands on the human operator, although some human intervention is still required to perform such tasks as insulator recognition, positional adjustments of the robot, and guidance toward electric lines and insulators. We develop a 3D object recognition and localization system for robot's positional adjustment. For localization to be useful and practical, the method is designed to be insensitive to outliers. Our algorithm is general enough to be applied not only to our dual-armed mobile robots but also to other teleoperation robots. This algorithm should greatly reduce the burden of human operators when it is applied to the robot. This paper first describes our algorithm, and then presents a performance evaluation.
Yasuyuki Someya, Kentaro Kawamura, Kiminori Hasegawa, Katsushi Ikeuchi
ICPR4
2000 Expanding Possible View Points of Virtual Environment Using Panoramic Images
abstract
Presents a method for creating a 3D virtual broad city environment with walk-through systems based on image-based rendering (lBR). In that virtual city, people can move rather freely and look at arbitrary views. The strength of our method is that we are able to easily render any view from an arbitrary point to an arbitrary direction on the ground in a virtual environment; previous methods, on the other hand, have strong restraints concerning their reconstructable areas. One of the other applications of our method is a driving simulator in the ITS domain. We can generate any view on any lane on the road from images taken by running along just one lane. Our method first captures panoramic images running along a straight line, indexing the capturing position of each image. The rendering process consists of selecting some suitable slits divided vertically from stored images, and reassembling them to create an image from a novel observation point.
Takuji Takahashi, Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
ICPR3
2000 Robust Localization for 3D Object Recognition Using Local EGI and 3D Template Matching with M-Estimators
abstract
A teleoperated system in a robot greatly reduces the demands on the human operator, although some human intervention is still required to perform such tasks as insulator recognition, positional adjustments of the robot, and guidance toward electric lines and insulators. In order to automate some of the robot's capabilities, we have developed a 3D object-localization method for the robot's positional adjustment. The method is designed to be insensitive to noise, outliers and occlusions while, at the same time, it has optimal run-time efficiency. The main contribution of our algorithm is the use of an objective function which is specified to reduce the effect of noise and outliers in the range image and a method for minimizing this function. The objective function is efficiently minimized by dynamically recomputing correspondences as the pose improves. Our algorithm is general enough to be applied not only to our dual-armed mobile robots but also to other teleoperation robots. This algorithm should greatly reduce the burden of operators when applied. This paper first describes our algorithm, and then presents a performance evaluation.
Kentaro Kawamura, Kiminori Hasegawa, Yasuyuki Someya, Yoichi Sato 0001, Katsushi Ikeuchi
ICRA5
2000 Symbolic Representation of Trajectories for Skill Generation
abstract
The completion of robot programs requires long development time and much effort. To shorten this programming time and minimize the effort, we have been developing a system which we refer to as "assembly-plan-from-observation (APO) system;" this system provides the ability for a robot to observe a human performing an assembly tasks, understand the tasks, and subsequently generate a program to perform that same task. One of the necessary tasks in APO is to create a trajectory of robot hand movement from observing human performance. The previous system developed a direct observation method based on the trajectory of a human movement. Though simple and handy, the system was susceptible to noise. This paper proposes a method to make the observation robust against noise by using symbolic representations of a trajectory based on contact analysis. The system divides the trajectory into small segments based on the contact analysis, then allocates an operation element referred to as a sub-skill to those segments; the result is a robust trajectory-based APO system.
Hirohisa Tominaga, Jun Takamatsu, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
ICRA5
2000 Recognition of human task by attention point analysis
abstract
This paper presents a novel method of constructing a human task model by attention point (AP) analysis. The AP analysis consists of two steps: at the first step, it broadly observes human task, constructs rough human task model and finds APs which require detailed analysis; and at the second step, by applying time-consuming analysis on APs in the same human task, it can enhance the human task model. This human task model is highly abstracted and is able to change the degree of abstraction adapting to the environment so as to be applicable in a different environment. We describe this method and its implementation using data gloves and a stereo vision system. We also show an experimental result in which a real robot observed a human task and performed the same human task successfully in a different environment using this model.
Koichi Ogawara, Soshi Iba, Tomikazu Tanuki, Hiroshi Kimura, Katsushi Ikeuchi
IROS5
2000 Extracting manipulation skills from observation
abstract
The completion of robot programs requires a long development time and much effort. To shorten this programming time and to minimize the effort, we have been developing a system which we refer to as the "Assembly Plan from Observation (APO) system". This system provides the ability for a robot to observe a human performing an assembly task, to recognize the task and to then generate a program to perform that same task. One of the necessary tasks in APO is to create the trajectory of a robot hand movement after having observed a human's performance. The previous system developed a direct observation method based on the trajectory of a human's movement. Although it was simple and handy, the system was susceptible to noise. This paper proposes a method to make the observation robust against noise by analyzing topological contact relations. The system divides the trajectory into small segments based on this contact analysis, and then allocates an operation element, referred to as a "sub-skill", to those segments. The result is a robust, trajectory-based APO system.
Jun Takamatsu, Hirohisa Tominaga, Koichi Ogawara, Hiroshi Kimura, Katsushi Ikeuchi
IROS5
2000 Appearance-based visual learning and object recognition with illumination invariance
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
Mach. Vis. Appl.3
2000 Special issue on vision applications and technology for intelligent vehicles: part I-infrastructure
Alberto Broggi, Katsushi Ikeuchi, Charles E. Thorpe
IEEE Trans. Intell. Transp. Syst.2
2000 Traffic monitoring and accident detection at intersections
abstract
We have developed an algorithm, referred to as spatio-temporal Markov random field, for traffic images at intersections. This algorithm models a tracking problem by determining the state of each pixel in an image and its transit, and how such states transit along both the x-y image axes as well as the time axes. Our algorithm is sufficiently robust to segment and track occluded vehicles at a high success rate of 93%-96%. This success has led to the development of an extendable robust event recognition system based on the hidden Markov model (HMM). The system learns various event behavior patterns of each vehicle in the HMM chains and then, using the output from the tracking system, identifies current event chains. The current system can recognize bumping, passing, and jamming. However, by including other event patterns in the training set, the system can be extended to recognize those other events, e.g., illegal U-turns or reckless driving. We have implemented this system, evaluated it using the tracking results, and demonstrated its effectiveness.
Shunsuke Kamijo, Yasuyuki Matsushita, Katsushi Ikeuchi, Masao Sakauchi
IEEE Trans. Intell. Transp. Syst.3
1999 Eigen-Texture Method: Appearance Compression Based on 3D Model
abstract
Image-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. Extensive research on these two methods has been made in CV and CG communities. However, both methods still have several drawbacks when it comes to applying them to the mixed reality where we integrate such virtual images with real background images. To overcome these difficulties, we propose a new method which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. The 3D model is generated from a sequence of range images. The Eigen-Texture method is practical because it does not require any detailed reflectance analysis of the object surface, and has great advantages due to the accurate 3D geometric models. This paper describes the method, and reports on its implementation.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR3
1999 Measurement of Surface Orientations of Transparent Objects Using Polarization in Highlight
abstract
This paper proposes a method for obtaining surface orientations of transparent objects using polarization in highlight. Since the highlight, the specular component of reflection light from objects, is observed only near the specular direction, it appears merely on limited parts on an object surface. In order to obtain orientations of a whole object surface, we employ a spherical extended light source. This paper reports its experimental apparatus, a shape recovery algorithm, and its performance evaluation.
Megumi Saito, Hiroshi Kashiwagi, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR4
1999 Illumination Distribution from Shadows
abstract
The image irradiance of a three-dimensional object is known to be the function of three components: the distribution of light sources, the shape, and reflectance of a real object surface. In the past, recovering the shape and reflectance of an object surface from the recorded image brightness has been intensively investigated. On the other hand, there has been little progress in recovering illumination from the knowledge of the shape and reflectance of a real object. In this paper, we propose a new method for estimating the illumination distribution of a real scene from image brightness observed on a real object surface in that scene. More specifically, we recover the illumination distribution of the scene from a radiance distribution inside shadows cast by an object of known shape onto another object surface of known shape and reflectance. By using the occlusion information of the incoming light, we are able to reliably estimate the illumination distribution of a real scene, even in a complex illumination environment.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR3
1999 Appearance Compression and Synthesis based on 3D Model for Mixed Reality
abstract
Rendering photorealistic virtual objects from their real images is one of the main research issues in mixed reality systems. We previously proposed the Eigen-Texture method (K. Nishino et al., 1999), a new rendering method for generating virtual images of objects from their real images to deal with the problems posed by past work in image based methods and model based methods. Eigen-Texture method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. However, we had a serious limitation in our system, due to the alignment problem of the 3D model and color images. We deal with this limitation by solving the alignment problem; we do this by using the method originally designed by P. Viola (1995). The paper describes the method and reports on how we implement it.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV3
1999 Illumination Distribution from Brightness in Shadows: Adaptive Estimation of Illumination Distribution with Unknown Reflectance Properties in Shadow Regions
abstract
An approach to extract watersheds and watercourses, as well as their corresponding valleys and hills, from images with subpixel precision is proposed. The critical points of the terrain are essential as the starting points for the construction of these separatrices. They are extracted efficiently with subpixel precision using an approach based on derivatives of Gaussian filters. The separatrices are extracted by integrating their defining differential equation. Finally, the hills and valleys are constructed by an efficient graph search algorithm. Examples show the quality of the results that can be achieved with the proposed approach.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV3
1999 Local-feature based vehicle recognition in infra-red images using parallel vision board
abstract
The paper describes a method for vehicle recognition, in particular, for recognizing a vehicle's make and model. Our system employs infra-red images so that we can use the same algorithm both day and night. Originally, the algorithm was the eigen-window method based on local features, but it has been changed to a vector quantization based algorithm which was originally proposed by J. Krumm (1997), to implement on an IMAP parallel image processing board. Any of these systems, based on both the eigen-window method and the vector quantization method, make a compressed database of local features for the algorithm of a target vehicle from given training images in advance; the system then matches a set of local features in the input image with those in training images for recognition. This method has the following three advantages: (1) it can detect even if part of the target vehicle is occluded; (2) it can detect even if the target vehicle is translated due to running out of lanes; (3) it does not require us to segment a vehicle from input images. The above advantages have been confirmed by performing outdoor experiments.
Masataka Kagesawa, Shinichi Ueno, Katsushi Ikeuchi, Hiroshi Kashiwagi
IROS3
1999 Task-model based human robot cooperation using vision
abstract
In order to assist a human, the robot must recognize human motion in real time by vision, and must plan and execute the needed assistance motion based on the task purpose and the context. In this research, we tried to solve such problems. We defined the abstract task model, analyzed the human demonstration by using events and an event stack, and automatically generated the task models needed in the assistance by the robot. The robot planned and executed the appropriate assistance motions based on the task: models according to the human motions in the cooperation with the human. We implemented a 3D object recognition system and a human grasp recognition system by using trinocular stereo color cameras and a real time range finder. The effectiveness of these methods was tested through an experiment in which the human and the robotic hand assembled toy parts in cooperation.
Hiroshi Kimura, Tomoyuki Horiuchi, Katsushi Ikeuchi
IROS3
1999 Automatic modeling of a 3D city map from real-world video
abstract
Mixed reality (MR) systems which integrate the virtual world and the real world have become a major topic in the research area of multimedia. As a practical application of these MR systems, we propose an efficient method for making a 3D map from real-world video data. The proposed method is an automatic organization method focusing on video objects to describe video data in an efficient way, i.e., by collating the real-world video data with map information using DP matching. To demonstrate the reliability of this method, we describe successful experiments that we performed using 3D information obtained from the real-world video data.
Hiroshi Kawasaki, Tomoyuki Yatabe, Katsushi Ikeuchi, Masao Sakauchi
ACM Multimedia (1)3
1999 Acquiring a Radiance Distribution to Superimpose Virtual Objects onto a Real Scene
abstract
This paper describes a new method for superimposing virtual objects with correct shadings onto an image of a real scene. Unlike the previously proposed methods, our method can measure a radiance distribution of a real scene automatically and use it for superimposing virtual objects appropriately onto a real scene. First, a geometric model of the scene is constructed from a pair of omnidirectional images by using an omnidirectional stereo algorithm. Then, radiance of the scene is computed from a sequence of omnidirectional images taken with different shutter speeds and mapped onto the constructed geometric model. The radiance distribution mapped onto the geometric model is used for rendering virtual objects superimposed onto the scene image. As a result, even for a complex radiance distribution, our method can superimpose virtual objects with convincing shadings and shadows cast onto the real scene. We successfully tested the proposed method by using real images to show its effectiveness.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Vis. Comput. Graph.3
1998 Recognition of Urban Scene Using Silhouette of Buildings and City Map Database
Wei Ku, Katsushi Ikeuchi, Masao Sakauchi
ACCV (2)3
1998 Appearance Based Visual Learning and Object Recognition with Illumination Invariance
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
ACCV (2)3
1998 Measuring Object Surface Shape and Reflectance Properties
Yoichi Sato 0001, Mark D. Wheeler, Katsushi Ikeuchi
ACCV (2)3
1998 3D Line's Extraction from 2D Spatio-temporal Image Created by Sine Slit
Pingtao Wang, Katsushi Ikeuchi, Masao Sakauchi
ACCV (2)2
1998 Consensus Surfaces for Modeling 3D Objects from Multiple Range Images
abstract
In this paper, we present a robust method for creating a triangulated surface mesh from multiple range images. Our method merges a set of range images into a volumetric implicit-surface representation which is converted to a surface mesh using a variant of the marching-cubes algorithm. Unlike previous techniques based on implicit-surface representations, our method estimates the signed distance to the object surface by finding a consensus of locally coherent observations of the surface. We call this method the consensus-surface algorithm. This algorithm effectively eliminates many of the troublesome effects of noise and extraneous surface observations without sacrificing the accuracy of the resulting surface. We utilize octrees to represent volumetric implicit surfaces-effectively reducing the computation and memory requirements of the volumetric representation without sacrificing accuracy of the resulting surface. We present results which demonstrate that our consensus-surface algorithm can construct accurate geometric models from rather noisy input range data.
Mark D. Wheeler, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV3
1998 Localization of insulators in electric distribution systems by using 3D template matching from multiple range images
abstract
Kyushu Electric has developed a dual-armed mobile-robot for use in electricity distribution systems. Although some human intervention is still required, the robot greatly reduces the demands on the human operator. In order to automate some of the robot's capabilities, we have developed a 3D object-localization method for robot positional adjustment. The method is designed to be insensitive to noise and outliers while, at the same time, it has optimal run-time efficiency. The paper first describes our algorithm, and then presents a performance evaluation.
Kentaro Kawamura, Mark D. Wheeler, Osamu Yamashita, Yoichi Sato 0001, Katsushi Ikeuchi
IROS5
1998 Task-Oriented Generation of Visual Sensing Strategies in Assembly Tasks
abstract
This paper describes a method of systematically generating visual sensing strategies based on knowledge of the assembly task to be performed. Since visual sensing is usually performed with limited resources, visual sensing strategies should be planned so that only necessary information is obtained efficiently. The generation of the appropriate visual sensing strategy entails knowing what information to extract, where to get it, and how to get it. This is facilitated by the knowledge of the task, which describes what objects are involved in the operation, and how they are assembled. In the proposed method, using the task analysis based on face contact relations between objects, necessary information for the current operation is first extracted. Then, visual features to be observed are determined using the knowledge of the sensor, which describes the relationship between a visual feature and information to be obtained. Finally, feasible visual sensing strategies are evaluated based on the predicted success probability, and the best strategy is selected. Our method has been implemented using a laser range finder as the sensor. Experimental results show the feasibility of the method, and point out the importance of task-oriented evaluation of visual sensing strategies.
Jun Miura, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 A quasi-linear method for computing and projecting onto c-surfaces: planar case
abstract
This paper presents a general method to compute configuration space (c-space) obstacle surfaces (c-surfaces) in planar quaternion space. We extend the method to find the projection of a given point in c-space onto the c-surface We parameterize the general c-surface using a rotation angle and the vector of translation parameters of the individual contacts. We first compute the domain of the rotation parameter. Then, we can setup the translation parameters in a linear equation. The solution of this equation using singular value decomposition gives us the exact parameters of translation. We can extend this quasi-linear method to project a point in c-space onto the c-surface. We implement our theory on the assembly plan from observation (APO) system. The APO observes discrete instants of an assembly task and reconstructs the compliant motion plan employed in the task. We compute the contacts at each observed instant and the corresponding c-surface. We then interpolate the path on each c-surface to obtain segments of the path. The complete motion plan will be the concatenation of the connected path segments.
George V. Paul, Katsushi Ikeuchi
ICRA2
1997 Visual learning and object verification with illumination invariance
abstract
This paper describes a method for recognizing partially occluded objects to realize a bin-picking task under different levels of illumination brightness by using the eigenspace analysis. In the proposed method, a measured color in the RGB color space is transformed into the HSV color space. Then, the hue of the measured color, which is invariant to change in illumination brightness and direction, is used for recognizing multiple objects under different levels of illumination conditions. The proposed method was applied to real images of multiple objects under different illumination conditions, and the objects were recognized and localized successfully.
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
IROS3
1997 A quasi-linear method for computing and projecting onto c-surfaces: general case
abstract
This paper presents a general method to compute configuration space (c-space) obstacle surfaces (c-surfaces) in dual quaternion space and for projecting points onto them. We parameterize the c-surface using the rotation angles of the object and the vector of translation parameters of the individual contacts. Once we compute the domain of the rotation parameters, we can setup the translation parameters in a linear equation. The singular value decomposition of this equation gives us with the exact parameters of translation. We extend the theory to find the projection of a point in c-space onto the c-surface. We implement our theory on the assembly plan from observation (APO) system. The APO observes discrete instants of an assembly task and reconstructs the compliant motion plan employed in the task. We compute the contacts at each observed instant and the corresponding c-surface. We then interpolate the path on each c-surface to obtain segments of the path. The complete motion plan will be the concatenation of the connected path segments.
George V. Paul, Katsushi Ikeuchi
IROS2
1997 Object shape and reflectance modeling from observation
abstract
An object model for computer graphics applications should contain two aspects of information: shape and reflectance properties of the object.A number of techniques have been developed for modeling object shapes by observing real objects.In contrast, attempts to model reflectance properties of real objects have been rather limited.In most cases, modeled reflectance properties are too simple or too complicated to be used for synthesizing realistic images of the object.In this paper, we propose a new method for modeling object reflectance properties, as well as object shapes, by observing real objects.First, an object surface shape is reconstructed by merging multiple range images of the object.By using the reconstructed object shape and a sequence of color images of the object, parameters of a reflection model are estimated in a robust manner.The key point of the proposed method is that, first, the diffuse and specular reflection components are separated from the color image sequence, and then, reflectance parameters of each reflection component are estimated separately.This approach enables estimation of reflectance properties of real objects whose surfaces show specularity as well as diffusely reflected lights.The recovered object shape and reflectance properties are then used for synthesizing object images with realistic shading effects under arbitrary illumination conditions.
Yoichi Sato 0001, Mark D. Wheeler, Katsushi Ikeuchi
SIGGRAPH3
1997 3D shape and reflectance morphing
abstract
The paper describes a new method for 3D shape and reflectance morphing of two real 3D objects. Our morphing method consists of two components: shape and reflectance property measurement, and smooth interpolation of those measured properties. Unlike other morphing techniques, the proposed method can create intermediate images with correct shading such as highlights and shadows.
Yoichi Sato 0001, Imari Sato, Katsushi Ikeuchi
Shape Modeling International3
1997 Detectability, Uniqueness, and Reliability of Eigen Windows for Stable Verification of Partially Occluded Objects
abstract
This paper describes a method for recognizing partially occluded objects for bin-picking tasks using eigenspace analysis, referred to as the "eigen window" method, that stores multiple partial appearances of an object in an eigenspace. Such partial appearances require a large amount of memory space. Three measurements, detectability, uniqueness, and reliability, on windows are developed to eliminate redundant windows and thereby reduce memory requirements. Using a pose clustering technique, the method determines the pose of an object and the object type itself. We have implemented the method and verified its validity.
Kohtaro Ohba, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 An Integral Approach to Free-Form Object Modeling
abstract
Presents an approach to free-form object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views, so that correspondence can be established among different views. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the techniques of principal component analysis with missing data (PCAMD). Experiments show that our approach is robust against noise and mismatch.
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 Toward automatic robot instruction from perception-mapping human grasps to manipulator grasps
abstract
Our approach of programming a robot is by direct human demonstration. The system observes a human performing the task, recognizes the human grasp, and maps it onto the manipulator. This paper describes how an observed human grasp can be mapped to that of a given general-purpose manipulator for task replication. Planning the manipulator grasp based upon the observed human grasp is done at two levels: the functional and physical levels. Initially, at the functional level, grasp mapping is achieved at the virtual finger level; the virtual finger is a group of fingers acting against an object surface in a similar manner. Subsequently, at the physical level, the geometric properties of the object and manipulator are considered in fine-tuning the manipulator grasp. Our work concentrates on power or enveloping grasps and the fingertip precision grasps. We conclude by showing an example of an entire programming cycle from human demonstration to robot execution.
Sing Bing Kang, Katsushi Ikeuchi
IEEE Trans. Robotics Autom.2
1996 Invariant histograms and deformable template matching for SAR target recognition
abstract
Recognizing a target in synthetic-aperture radar (SAR) images is an important, yet challenging, application of the model-based vision technique. This paper describes a model-based SAR recognition system based on invariant histograms and deformable template matching techniques. An invariant histogram is a histogram of invariant values defined by geometric features such as points and lines in SAR images. Although a few invariants are sufficient to recognize a target, we use a histogram of all invariant values given by all possible target feature pairs. This redundant histogram enables robust recognition under severe occlusions typical in SAR recognition scenarios. Multi-step deformable template matching examines the existence of an object by superimposing templates over potential energy field generated from images or primitive features. It determines the template configuration which has the minimum deformation and the best alignment of the template with features. The deformability of the template absorbs the instability of SAR features. We have implemented the system and evaluated the system performance using hybrid SAR images, generated from synthesized model signatures and real SAR background signatures.
Katsushi Ikeuchi, Takeshi Shakunaga, Mark D. Wheeler, Taku Yamazaki
CVPR1
1996 On 3D Shape Similarity
abstract
This paper addresses the problem of 3D shape similarity between closed surfaces. A curved or polyhedral 3D object of genus zero is represented by a mesh that has nearly uniform distribution with known connectivity among mesh nodes. A shape similarity metric is defined based on the L/sub 2/ distance between the local curvature distributions over the mesh representations of the two objects. For both convex and concave objects, the shape metric can be computed in time O(n/sup 2/), where n is the number of tessellations of the sphere or the number of meshes which approximate the surface. Experiments show that our method produces good shape similarity measurements.
Harry Shum, Martial Hebert, Katsushi Ikeuchi
CVPR3
1996 Recognition of the multi-specularity objects using the eigen-window
abstract
This paper describes a method for recognizing partially occluded specularity objects for bin-picking tasks using the eigen-space analysis. Although effective in recognizing an isolated object, as was shown by Murase and Nayar, the current method can not be applied to partially occluded objects that are typical in bin-picking tasks. The analysis also requires that the object is centered in an image before recognition. These limitations of the eigen-space analysis are due to the fact that the whole appearance of an object is utilized as a template for the analysis. We propose a new method, referred to as the "eigen-window" method, that stores multiple partial appearances of an object in the eigen-space. Such partial appearances require a large number of memory space. To reduce the memory requirement by avoiding redundant windows and to select only effective windows to be stored, a similarity measure among windows is developed. Using a pose clustering method among windows, the method determines the pose of an object. We have implemented the method and verify the validity of the method.
Kohtaro Ohba, Katsushi Ikeuchi
ICPR2
1996 Hand action perception for robot programming
abstract
This paper presents a general and robust approach to hand action perception for automatic robot programming using depth image sequences. The human instructor must simply demonstrate an assembly task in front of a vision system in the human world; no dataglove or special markings are necessary. The recorded image sequences are used to recover a depth image sequence for model-based human hand and object tracking to form the perceptual data stream. The data stream is then segmented and interpreted for generating a task sequence which describes the human hand action and the relationship between the manipulated object and the hand. The task sequence might be composed of a series of subtasks and each subtask involves four phases: approaching, pre-manipulating, manipulating and departing. In this paper we also discuss a robot system that replicates the observed task and automatically validates the replication results in the robot world.
Yunde Jiar, Mark D. Wheeler, Katsushi Ikeuchi
IROS3
1996 Recognition of the multi specularity objects for bin-picking task
abstract
This paper describes a method for recognizing partially occluded objects for bin-picking tasks using the eigen-space analysis. Although effective in recognizing an isolated object, as was shown by Murase and Nayar (1995), the current method can not be applied to piratically occluded objects that are typical in bin-picking tasks. The analysis also requires that the object is centered in an image before recognition. These limitations of the eigen-space analysis are due to the fact that the whole appearance of an object is utilized as a template for the analysis. We propose a new method, referred to as the "eigen-window" method, that stores multiple partial appearances of an object in the eigen-space. Such partial appearances require a large number of memory space. To reduce the memory requirement by avoiding redundant windows and to select only effective windows to be stored, a similarity measure among windows is developed. Using a pose clustering method among windows, the method determines the pose of an object and the object type of itself. We have implemented the method and verify the validity of the method.
Kohtaro Ohba, Katsushi Ikeuchi
IROS2
1996 Modeling planar assembly paths from observation
abstract
This paper describes a system for obtaining the motion plan for a planar assembly task, given a sequence of observations of a human performing the task. The motion plan in configuration space is a series of connected path segments lying outside and on the configuration space obstacle. We use the observed configurations of the assembled objects to selectively compute the features of the c-space obstacle on which the path lies. We project the observed configurations onto these features and reconstruct the path segments. The connected path segments form the model of the observed task and can be used to program a robot to repeat the task. We demonstrate the system using the planar peg in hole task.
George V. Paul, Katsushi Ikeuchi
IROS2
1996 Reflectance Analysis for 3D Computer Graphics Model Generation
Yoichi Sato 0001, Katsushi Ikeuchi
CVGIP Graph. Model. Image Process.2
1996 Extracting the Shape and Roughness of Specular Lobe Objects Using Four Light Photometric Stereo
abstract
We propose a noncontact method for the measurement of surface shape and surface roughness. The method, which we call "four light photometric stereo", uses four lights, which sequentially illuminate the object under inspection, and a video camera for taking images of the object. The method has successfully been applied to a number of real objects.
Fredric Solomon, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 Iterative Smoothed Residuals: A Low-Pass Filter for Smoothing With Controlled Shrinkage
abstract
We present a linear smoothing operator which has low-pass characteristics similar to a Butterworth filter and limited spatial extent similar to a Gaussian. The smoothing operator also has closed forms in the spatial and frequency domains which facilitate analysis and implementation. A formula is derived that allows us to explicitly control shrinkage.
Mark D. Wheeler, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 A Robot System that Observes and Replicates Grasping Tasks
abstract
To alleviate the problem of overwhelming complexity in grasp synthesis and path planning associated with robot task planning, we adopt the approach of teaching the robot by demonstrating in front of it. The system has four components: the observation system, the grasping task recognition module, the task translator and the robot system. The observation system comprises an active multibaseline stereo system and a dataglove. The data stream recorded is then used to track object motion; this paper illustrates how complimentary sensory data can be used for this purpose. The data stream is also interpreted by the grasping task recognition module, which produces higher levels of abstraction to describe both the motion and actions taken in the task. The resulting information are provided to the task translator which creates commands for the robot system to replicate the observed task. In this paper we describe how these components work with special emphasis on the observation system. The robot system that we use to perform the grasping tasks comprises the PUMA 560 arm and the Utah/MIT hand.>
Sing Bing Kang, Katsushi Ikeuchi
ICCV2
1995 Task-Oriented Generation of Visual Sensing Strategies
abstract
In vision-guided robotic operations, vision is used for extracting necessary information for achieving the task. Since visual sensing is usually performed with limited resources, visual sensing strategies should be planned so that only necessary information is obtained efficiently. This paper describes a method of systematically generating visual sensing strategies based on knowledge of the task to be performed. The generation of the appropriate visual sensing strategy entails knowing what information to extract, where to get it, and how to get it. This is facilitated by the knowledge of the task, which describes what objects are involved in the operation, and how they are assembled. Our method has been implemented using a laser range finder as the sensor. Experimental results show the feasibility of the method, and point out the importance of task-oriented evaluation of visual sensing strategies.>
Jun Miura, Katsushi Ikeuchi
ICCV2
1995 An Integral Approach to Free-Formed Object Modeling
abstract
Presents a new approach to free-formed object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views so that correspondences can be established among different views. The global resampling algorithm is based on the spherical attribute image (SAI) previously introduced in the context of object representation and recognition. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the technique of principal component analysis with missing data (PCAMD). Experiments using real range images show that our approach is robust against noise and mismatches, and generates accurate object models.>
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
ICCV3
1995 Generating Visual Sensing Strategies in Assembly Tasks
abstract
It is generally very difficult, if not impossible, for a robot to perform fine manipulation tasks without the benefit of some form of sensory feedback during actual task execution. As a result, sensing planning is an important component in assembly task planning. This paper describes a method of generating visual sensing strategies based on knowledge of the task to be performed. The generation of the appropriate visual sensing strategy entails knowing what information to extract and where to get it. This is facilitated by the knowledge of the task, which describes how objects are assembled. This knowledge, coupled with known sensor modeling, results in an abstract template of sensing strategy called the sensing task model. By instantiating the appropriate sensing task model at planning time, the sensing strategy is efficiently generated. Our method has been implemented using a laser range finder as the sensor. Experimental results involving typical assembly tasks show the feasibility of the method.
Jun Miura, Katsushi Ikeuchi
ICRA2
1995 Partitioning COntact-State Space Using The Theory of Polyhedral Convex Cones
abstract
The assembly plan from observation (APO) system observes a human operator perform an assembly task, analyzes the observations, models the task and generates the programs for the robot to perform the same task. A major component of the APO system is the task recognition module, which models the observed task. The task model in the APO context is defined as a sequence of assembly states of the part being assembled and the actions which cause the transition between the states. The state of the assembled part is based on its freedom, which can be computed from the geometry of the contacts between the part and its environment. This freedom can be represented as a polyhedral convex cone (PCC) in screw space. We show that any contact configuration can be classed into a finite number of contact states. These contact states correspond to typologically distinct intersections of the PCC with a linear subspace T in screw space. The models of any observed task can be represented as a path in a transition graph obtained from these contact states. We illustrate the theory by implementing the APO system for polygonal objects assembled in a plane using rotation and translation.
George V. Paul, Katsushi Ikeuchi
ICRA2
1995 An Illumination Planner for Lambertian polyhedral Objects
abstract
The measurement of shape is a basic object inspection task. We use a noncontact method to determine shape called photometric stereo. The method uses three light sources which sequentially illuminate the object under inspection and a video camera for taking intensity images of the object. A significant problem with using photometric stereo is determining where to place the 3 light sources and the video camera. In order to solve this problem, we have developed an illumination planner that determines how to position the three light sources and the video camera around the object. The planner determines how to position light sources around an object so that we illuminate a specified set of faces in an efficient manner and so that we obtain an accurate measurement. From a high level, our planner has three major inputs: the CAD model of the object to be inspected, a noise model for our sensor, and a reflectance model for the object to be inspected. We have experimentally verified that the plans generated by the planner are valid and accurate.
Fredric Solomon, Katsushi Ikeuchi
ICRA2
1995 Assembly of flexible objects without analytical models
abstract
The ability of manipulating flexible objects, such as rubber belts and paper sheets, is important in automated manufacturing systems. This paper describes a novel approach to assembly of flexible objects. The operation dealt with in this paper is to assemble a rubber belt with fixed pulleys. By analyzing possible states of the belt based on the empirical knowledge of the belt, one can derive a method to have not only the action planning but also the visual verification planning. The authors have implemented a belt assembly system using two manipulators and a laser range finder as the sensor, and succeeded in performing the belt-pulley assembly. Extension of the authors' approach to other kinds of assembly of flexible objects is also discussed.
Jun Miura, Katsushi Ikeuchi
IROS (2)2
1995 Modelling planar assembly tasks: representation and recognition
abstract
The assembly plan from observation (APO) system observes a human operator perform an assembly task, analyzes the observations, models the task, and generates the programs for the robot to perform the same task. The task model of the observed task is defined as a sequence of contact states of the part being assembled and the motion which causes the transition between the states. The freedom of the assembled part can be represented as a polyhedral convex cone (PCC) in screw space, which can belong to one of a finite number of distinct contact states. These contact states correspond to topologically distinct intersections of PCCs with a linear subspace T in screw space. Any observed assembly task can be represented as a finite sequence of critical contact states and the motion between them. The abstract task model is used to program the robot to execute the observed assembly task. We illustrate the application of the theory by implementing the APO system for assemblies in a plane.
George V. Paul, Katsushi Ikeuchi
IROS (1)2
1995 Building 3-D Models from Unregistered Range Images
Ken Higuchi, Martial Hebert, Katsushi Ikeuchi
CVGIP Graph. Model. Image Process.3
1995 Recent Progress in CAD-Based Vision
Katsushi Ikeuchi, Patrick J. Flynn
Comput. Vis. Image Underst.1
1995 A Spherical Representation for Recognition of Free-Form Surfaces
abstract
Introduces a new surface representation for recognizing curved objects. The authors approach begins by representing an object by a discrete mesh of points built from range data or from a geometric model of the object. The mesh is computed from the data by deforming a standard shaped mesh, for example, an ellipsoid, until it fits the surface of the object. The authors define local regularity constraints that the mesh must satisfy. The authors then define a canonical mapping between the mesh describing the object and a standard spherical mesh. A surface curvature index that is pose-invariant is stored at every node of the mesh. The authors use this object representation for recognition by comparing the spherical model of a reference object with the model extracted from a new observed scene. The authors show how the similarity between reference model and observed data can be evaluated and they show how the pose of the reference object in the observed scene can be easily computed using this representation. The authors present results on real range images which show that this approach to modelling and recognizing 3D objects has three main advantages: (1) it is applicable to complex curved surfaces that cannot be handled by conventional techniques; (2) it reduces the recognition problem to the computation of similarity between spherical distributions; in particular, the recognition algorithm does not require any combinatorial search; and (3) even though it is based on a spherical mapping, the approach can handle occlusions and partial views.>
Martial Hebert, Katsushi Ikeuchi, Hervé Delingette
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Principal Component Analysis with Missing Data and Its Application to Polyhedral Object Modeling
abstract
Observation-based object modeling often requires integration of shape descriptions from different views. To overcome the problems of errors and their accumulation, we have developed a weighted least-squares (WLS) approach which simultaneously recovers object shape and transformation among different views without recovering interframe motion. We show that object modeling from a range image sequence is a problem of principal component analysis with missing data (PCAMD), which can be generalized as a WLS minimization problem. An efficient algorithm is devised. After we have segmented planar surface regions in each view and tracked them over the image sequence, we construct a normal measurement matrix of surface normals, and a distance measurement matrix of normal distances to the origin for all visible regions over the whole sequence of views, respectively. These two matrices, which have many missing elements due to noise, occlusion, and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. A two-step algorithm is presented. After surface equations are extracted, spatial connectivity among the surfaces is established to enable the polyhedral object model to be constructed. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatching and generates accurate polyhedral object models.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Sensor Modeling, Probabilistic Hypothesis Generation, and Robust Localization for Object Recognition
abstract
In an effort to make object recognition efficient and accurate enough for real applications; we have developed three probabilistic techniques-sensor modeling, probabilistic hypothesis generation, and robust localization-which form the basis of a promising paradigm for object recognition. Our techniques effectively exploit prior knowledge to reduce the number of hypotheses that must be tested during recognition. Our recognition approach utilizes statistical constraints on the matches between image and model features. These statistical constraints are computed using a model of the entire sensing process-resulting in more realistic and tighter constraints on matches. The candidate hypotheses are pruned by probabilistic constraint satisfaction to select likely matches based on the image evidence and prior statistical constraints. The resulting hypotheses are ordered most-likely first for verification. Thus minimizing unnecessary verifications. The reliability of the verification decision is significantly increased by the use of a robust localization algorithm.>
Mark D. Wheeler, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Toward automatic robot instruction from perception-temporal segmentation of tasks from human hand motion
abstract
This paper describes work on the temporal segmentation of grasping task sequences based on human hand motion. The segmentation process results in the identification of motion breakpoints separating the different constituent phases of the grasping task. A grasping task is composed of three basic phases: pregrasp phase, static grasp phase, and manipulation phase. We show that by analyzing the fingertip polygon area (which is an indication of the hand preshape) and the speed of hand movement (which is an indication of the hand transportation), we can divide a task into meaningful action segments such as approach object (which corresponds to the pregrasp phase), grasp object, manipulate object, place object, and depart (a special case of the pregrasp phase which signals the termination of the task). We introduce a measure called the volume sweep rate, which is the product of the fingertip polygon area and the hand speed. The profile of this measure is also used in the determination of the task breakpoints.>
Sing Bing Kang, Katsushi Ikeuchi
IEEE Trans. Robotics Autom.2
1994 Object representation for object recognition
abstract
This paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.>
Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman
CVPR7
1994 Principal component analysis with missing data and its application to object modeling
abstract
Observation-based modeling can reduce the cost and effort of model constructions for tasks such as virtual reality environment. Object modeling from a sequence of range images has been formulated as a problem of principal component analysis with missing data (PCAMD), which can be generalized as a weighted least square (WLS) minimization problem. After all visible regions appeared over the whole sequence are segmented and tracked, a normal measurement matrix of surface normals and a distance measurement matrix of normal distances to the origin are constructed respectively. These two measurement matrices, with possibly many missing elements due to occlusion and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. The solution to the first WLS problem, which employs the quaternion representation of the rotation matrix, yields surface normals and rotation matrices. Subsequently the normal distances and translation vectors are computed by solving the second WLS problem. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatch because it produces a statistically optimal object model by making use of redundancy from multiple views. A toy house model from a sequence of real range images is presented.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
CVPR2
1994 Building 3-D Models from Unregistered Range Images
abstract
The authors describe an approach to building a three-dimensional model from a set of range images. The authors' goal is to build models of free-form surfaces obtained from arbitrary viewing directions, with no initial estimate of the relative viewing directions. The approach is based on building discrete meshes representing the surfaces observed in each of the range images, to map each of the meshes to a spherical image, and to compute the transformations between the views by matching the spherical images. The meshes are built using an iterative fitting algorithm previously developed; the spherical images are built by matching the nodes of the surface meshes to the nodes of a reference mesh on the unit sphere and by storing a measure of curvature at every node. The authors describe the algorithms used for building such models from range images and for matching them. The authors give results obtained using range images of complex objects.>
Ken Higuchi, Martial Hebert, Katsushi Ikeuchi
ICRA3
1994 Determination of Motion Breakpoints in a Task Sequence from Human Hand Motion
abstract
This paper describes the authors' work on the temporal segmentation of grasping task sequences based on human hand motion. The segmentation process results in the identification of motion breakpoints separating the different constituent phases of the grasping task. A grasping task is composed of three basic phases: pregrasp phase, static grasp phase, and manipulation phase. The authors show that by analyzing the fingertip polygon (preshape) area and the speed of hand movement, they can divide a task into meaningful action segments such as approach object, grasp object, manipulate object, place object, and depart. The authors introduce a measure called the volume sweep rate, which is the product of the fingertip polygon area and the hand speed. The profile of this measure is also used in the determination of the task breakpoints. The temporal task segmentation process is important as it serves as a preprocessing step to the characterization of the task phases. Once the breakpoints have been identified, further analyses such as grasp recognition and object motion extraction can then be carried out.>
Sing Bing Kang, Katsushi Ikeuchi
ICRA2
1994 Grasp Recoguition and Manipulative Motion Characterization from Human Hand Motion Sequences
abstract
We are developing a system capable of observing a human performing a task and understanding the task well enough to replicate it. This approach is called Assembly Plan from Observation. In order to replicate the observed task, we have to analyze the entire sequence. This can be done by first segmenting the task sequence into its constituent pre-grasp, grasp, and manipulation phases. This paper describes the different analyses that can be done subsequent to the temporal segmentation. These include human grasp recognition, extraction of object motion, and the spatiofrequency (spectrogram) analysis of the manipulation phase.>
Sing Bing Kang, Katsushi Ikeuchi
ICRA2
1994 Robot task programming by human demonstration: mapping human grasps to manipulator grasps
abstract
To alleviate the problem of overwhelming complexity in grasp synthesis and path planning associated with robot task planning, we adopt the approach of teaching the robot by demonstrating in front of it. A system with this programming technique is able to temporally segment a task into separate and meaningful parts for further individual analysis and recognize the human grasp employed in the task. With such derived information, this system would then map the human grasp to that of the given manipulator plan its trajectory, and proceed to execute the task. This paper describes how grasp mapping can be accomplished in our system. The mapping process essentially comprises three steps. The first step is local functional mapping, in which grasps of functionally equivalent fingers are established. This is followed by gross physical mapping which produces a kinematically feasible manipulator grasp. Finally, by carrying out local grasp adjustment using some task-related criterion, we arrive at a locally optimal manipulator grasp. We describe these steps in detail in this paper and show results of example grasp mappings.>
Sing Bing Kang, Katsushi Ikeuchi
IROS2
1994 Virtual reality modeling from a sequence of range images
abstract
Virtual reality object modeling from a sequence of range images has been formulated as a problem of principal component analysis with missing data (PCAMD), which can be generalized as a weighted least square (WLS) minimization problem. An efficient algorithm has been devised to solve the problem of PCAMD. After all visible P regions appeared over the whole sequence of F views are segmented and tracked, a 3F/spl times/P normal measurement matrix of surface normals and an F/spl times/P distance measurement matrix of normal distances to the origin are constructed respectively. These two measurement matrices, with possibly many missing elements due to occlusion and mismatching, enable us to formulate multiple view merging as a combination of two WLS problems. By combining information at both the signal level and the algebraic level, a modified Jarvis' march algorithm is proposed to recover the spatial connectivity among all the reconstructed surface patches. Experiments using synthetic data and real range images show that our approach is robust against noise and mismatch. A toy house model from a sequence of real range images is presented.>
Harry Shum, Katsushi Ikeuchi, Raj Reddy
IROS2
1994 Planning multiple observations for object recognition
Keith D. Gremban, Katsushi Ikeuchi
Int. J. Comput. Vis.2
1994 Toward an assembly plan from observation. I. Task recognition with polyhedral objects
abstract
The authors present the assembly-plan-from-observation (APO) method for robot programming. The APO method aims to build a system that has the capability of observing a human performing an assembly task, understanding the task based on the observation, and subsequently generating a robot program to achieve the same task. This paper focuses on the task recognition module (TRM), the main component of a complete APO system. The TRM recognizes object configurations before and after an assembly task, detects a configuration transition, and infers the assembly task that causes such a configuration transition. We assume that each assembly task aims to achieve a face contact relation between an object that has just been manipulated and the stationary environmental objects. We prepare abstract task models that associate transitions of face contact relations with assembly tasks that achieve such transitions. Next, we implement TRM in order to verify two issues: 1) that such a contact transition can be recovered from the output of the object recognizer; and 2) that given these relation transitions, it is possible to use the abstract task models to generate robot motion commands.>
Katsushi Ikeuchi, Takashi Suehiro
IEEE Trans. Robotics Autom.1
1993 Roughness and shape of specular lobe surfaces using photometric sampling method
abstract
An algorithm is proposed to determine surface orientation and roughness for specular lobe dominant surfaces under photometric sampling. From the image sequence, surface reflectance and orientation are obtained by determining the parameters of the reflectance model. The validity of this approach is demonstrated by applying it to real specular lobe dominant surfaces and examining orientation errors.>
Tesuo Kiuchi, Katsushi Ikeuchi
CVPR2
1993 Temporal-color space analysis of reflection
abstract
A method to analyze a sequence of color images is proposed. A series of images is examined in a four-dimensional space, which is called the temporal-color space, the axes of which are the three color axes (RGB) and one temporal axis. The significance of the temporal-color space lies in its ability to represent the change of image color with time. A conventional color space analysis yields the histogram of the colors in an image only at an instance of time. Conceptually, the two reflection components from the dichromatic reflection model, the specular reflection component and the body reflection component, form two subspaces in the temporal-color space. These two components can be extracted by principal component analysis. Using this fact, real color images are analyzed, and the two reflection components are separated successfully.>
Yoichi Sato 0001, Katsushi Ikeuchi
CVPR2
1993 A spherical representation for the recognition of curved objects
abstract
The authors introduce a surface representation for recognizing curved objects. The approach begins by representing an object by a discrete mesh of points built from range data or from a geometric model of the object. The mesh is computed from the data by deforming a standard shaped mesh, for example, an ellipsoid, until it fits the surface of the object. Local regularity constraints that the mesh must satisfy are defined. A canonical mapping is then defined between the mesh describing the object and a standard spherical mesh. A surface curvature index which is pose-invariant is stored at every node of the mesh. This object representation is used for recognition by comparing the spherical model of a reference object with the model extracted from a new observed scene. It is shown that the similarity between reference model and observed data can be evaluated, and it is also demonstrated that the pose of the reference object in the observed scene can be easily computed using this representation.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
ICCV3
1993 Toward assembly plan from observation - Task recognition with planar, curved and mechanical contacts
abstract
The authors have been developing a novel method for programming a robot, called the assembly-plan-from-observation (APO) method. The APO method aims to build a system that has threefold capabilities. It observes a human performing an assembly task, it understands the task based on this observation, and it generates a robot program to achieve the same task. This paper concentrates on the APO's main loop, which is task recognition. Using object recognition results, the task recognition module determines what kind of assembly task is performed. A previous system, recognizes assembly tasks which only handle polyhedral objects. The system reported here handles curved objects and other mechanical contacts as well. The authors define task models for these cases and show that task models are useful in recognizing assembly tasks, and that it is possible to generate robot motion commands for repeating the same assembly task.
Katsushi Ikeuchi, Masato Kawade, Takashi Suehiro
IROS1
1993 A grasp abstraction hierarchy for recognition of grasping tasks from observation
abstract
This work focuses on the abstraction hierarchy for a grasp which has been recognized from low-level hand-object interaction data. Previous work done on grasp classification and recognition is discussed. The proposed abstraction hierarchy is presented with illustrations, implementation issues, as well as experimental results. Issues pertaining to the conceptual analysis of the other aspects of recognizing grasping tasks are also presented. The authors report on the current status of the project and future work.
Sing Bing Kang, Katsushi Ikeuchi
IROS2
1993 Comment on "Numerical Shape from Shading and Occluding Boundaries"
Katsushi Ikeuchi
Artif. Intell.1
1993 The Complex EGI: A New Representation for 3-D Pose Determination
abstract
The complex extended Gaussian image (CEGI), a 3D object representation that can be used to determine the pose of an object, is described. In this representation, the weight associated with each outward surface normal is a complex weight. The normal distance of the surface from the predefined origin is encoded as the phase of the weight, whereas the magnitude of the weight is the visible area of the surface. This approach decouples the orientation and translation determination into two distinct least-squares problems. The justification for using such a scheme is twofold: it not only allows the pose of the object to be extracted, but it also distinguishes a convex object from a nonconvex object having the same EGI representation. The CEGI scheme has the advantage of not requiring explicit spatial object-model surface correspondence in determining object orientation and translation. Experiments involving synthetic data of two polyhedral and two smooth objects are presented to illustrate the feasibility of this method.>
Sing Bing Kang, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 Toward automatic robot instruction from perception-recognizing a grasp from observation
abstract
Deals with the programming of robots to perform grasping tasks. To do this, the assembly plan from observation (APO) paradigm is adopted, where the key idea is to enable a system to observe a human performing a grasping task, understand it, and perform the task with minimal human intervention. A grasping task is composed of three phases: pregrasp phase, static grasp phase, and manipulation phase. The first step in recognizing a grasping task is identifying the grasp itself. The proposed strategy of identifying the grasp is to map the low-level hand configuration to increasingly more abstract grasp descriptions. To achieve the mapping, a grasp representation is introduced, called the contact web, which is composed of a pattern of effective contact points between the hand and the object. A grasp taxonomy based on the contact web is also proposed as a tool to systematically identify a grasp. The grasp can be described at higher conceptual levels using a certain mapping function that results in an index called the grasp cohesive index. This index can be used to identify the grasp. Results from grasping experiments show that it is possible to distinguish between various types of grasps using the proposed contact web, grasp taxonomy and grasp cohesive index.>
Sing Bing Kang, Katsushi Ikeuchi
IEEE Trans. Robotics Autom.2
1992 Recognizing assembly tasks using face-contact relations
abstract
A novel method for programming a robot, called the assembly-plan-from-observation (APO) method, is proposed. The APO method aims to build a system that has the capability of observing a human performing an assembly task, understanding the task based on the observation, and generating the robot program to achieve the same task. Assembly relations that serve as the basic representation of each assembly task are defined. It is verified that such assembly relations can be recovered from the observation of human assembly tasks, and that from such assembly relations it is possible to generate robot motion commands to repeat the same assembly task. An APO system based on the assembly relations is demonstrated.>
Katsushi Ikeuchi, Takashi Suehiro
CVPR1
1992 Extracting the shape and roughness of specular lobe objects using four light photometric stereo
abstract
A noncontact method of measuring surface shape and surface roughness for part inspection is proposed. The method, called four-light photometric stereo, uses four lights that sequentially illuminate the object under inspection, and a video camera that takes images of the object. Conceptually, the problem has three parts: shape extraction, pixel segmentation, and roughness extraction. The shape information is produced directly by three-light and four-light photometric stereo methods. After shape information is obtained, statistical segmentation techniques can be applied to determine which pixels are specular and which are nonspecular. Then the specular pixels and shape information can be used, in conjunction with a simplified Torrance-Sparrow reflectance model, to determine the surface roughness. The method has successfully been applied to a number of synthetic and real objects. >
Fredric Solomon, Katsushi Ikeuchi
CVPR2
1992 Towards an assembly plan from observation. I. Assembly task recognition using face-contact relations (polyhedral objects)
abstract
The authors propose a novel method to program a robot, the assembly-plan-from-observation (APO) method. The APO method aims to build a system that has the capability of observing a human performing an assembly task, understanding the task based on the observation, and generating the robot program to achieve the same task. Assembly relations which serve as the basic representation of each assembly task are defined. It was verified that such assembly relations can be recovered from the observation of human assembly tasks, and that from such assembly relations, it is possible to generate robot motion commands to repeat the same assembly task. An APO system based on the assembly relations was demonstrated.>
Katsushi Ikeuchi, Takashi Suehiro
ICRA1
1992 Inspecting specular lobe objects using four light sources
abstract
The authors propose a noncontact method of measuring surface shape and surface roughness. The method uses four lights which sequentially illuminate the object under inspection, and a video camera for taking images of the object. Conceptually, the problem has three parts: shape extraction, pixel segmentation, and roughness extraction. The shape information is produced directly by three-light and four-light photometric stereo methods. After shape information is obtained, statistical segmentation techniques can be applied to determine which pixels are specular and which are nonspecular. Then, the specular pixels and shape information, and a simplified Torrance-Sparrow reflectance model can be used to determine the surface roughness. The method has successfully been applied to a number of synthetic and real objects.>
Fredric Solomon, Katsushi Ikeuchi
ICRA2
1992 Task Oriented Vision
Katsushi Ikeuchi, Martial Hebert
IROS1
1992 Grasp Recognition Using The Contact Web
abstract
We propose an approach to teach robots to ger- form grasping tasks. This approach is based on the Assembly Plan from Observation (APO) paradigm, where the key idea is to enable a system to observe a human performing a grasping task, understand it, and perform the task with minimal human intervention. A grasping task is composed of three phases: pre-grasp phase, static grasp phase, and manipulation phase. The first step in recognizing a grasping task is to identify the grasp itself (within the static grasp phase). We propose to identify the grasp by means of a grasp representation called the contact web which is composed of a pattern of effective contact points between the hand and the object. We also propose a grasp taxonomy based on the contact web to systematically identify a grasp. Results from grasping experiments show that it is possi- ble to distinguish between various types of grasps using the proposed contact web and grasp taxonomy.
Sing Bing Kang, Katsushi Ikeuchi
IROS2
1992 Towards An Assembly Plan From Observation: Part II: Correction Of Motion parameters Based On Fact Contact Constraints
abstract
We have been developing a novel method to program a robot, an APO (assembly-plan- from-observation) method. A human performs assem- bly operations in front of the APO system's TV caimera. The APO system recognizes such assembly operations and generates an assembly plan to repeat the :assembly operations using its robot arm. We have developped a APO system which extracts two kinds of information from observed object config- urations: 1) face-contact relation and 2) motion pa- rameters necessary to move objects around. Under the presence of error, due to error-contaminated motion pa- rameters, the system may fail to perform an assembly operation, although a face-contact relation and thus an assembly operation is correctly recovered. This paper proposes a method to correct eirr~lneous motion parameters based on a face contact re)(ation. We assume that a given face contact relation reflects the actual face contact correctly. Motion paramleters are determined by solving face contact constraint equaitions, given by face contact relations, simultaneously )using the singular value decomposition method. We implement this method in the APO system, apply the method to several assembly examples, and verify the effectiveness of tlhe method. are the most promising. Yet, these methods are often in- convepient and impractical.
Takashi Suehiro, Katsushi Ikeuchi
IROS2
1992 Why aspect graphs are not (yet) practical for computer vision
Olivier D. Faugeras, Joseph L. Mundy, Narendra Ahuja, Charles R. Dyer, Alex Pentland, Ramesh Jain 0001, Katsushi Ikeuchi, Kevin W. Bowyer
CVGIP Image Underst.7
1992 Model based recognition of specular objects using sensor models
Kosuke Sato, Katsushi Ikeuchi, Takeo Kanade
CVGIP Image Underst.2
1992 Shape representation and image segmentation using deformable surfaces
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
Image Vis. Comput.3
1991 Shape representation and image segmentation using deformable surfaces
abstract
A technique for constructing shape representation from images using free-form deformable surfaces is presented. The authors model an object as a closed surface that is deformed subject to attractive fields generated by input data points and features. Features affect the global shape of the surface, while data points control its local shape. This approach is used to segment objects even in cluttered or unstructured environments. The algorithm is general in that it makes few assumptions on the type of features, the nature of the data, and the type of objects. Results for a wide range of applications are presented: reconstruction of smooth isolated objects such as human faces, reconstruction of structured objects such as polyhedra, and segmentation of complex scenes with mutually occluding objects. The algorithm has been successfully tested using data from different sensors including grey-coding range finders and video cameras, using one or several images.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
CVPR3
1991 Determining 3-D object pose using the complex extended Gaussian image
abstract
A method based on the extended Gaussian image (EGI) which can be used to determine the pose of a 3-D object is presented. In this scheme, the weight associated with each outward surface normal is a complex weight. The normal distance of the surface from the predefined origin is encoded as the phase of the weight, while the magnitude of the weight is the visible area of the surface. This approach decouples the orientation and translation determination into two distinct least-squares problems. Experiments involving synthetic data of two polyhedral and two smooth objects as well as real range data of the same smooth objects indicate the feasibility of this method.>
Sing Bing Kang, Katsushi Ikeuchi
CVPR2
1991 A three-finger gripper for manipulation in unstructured environments
abstract
A gripper is described for manipulation in natural, unstructured environments. The specific manipulation task is to pick up surface material such as pebbles or small rocks in a natural terrain. The application is to give autonomous sampling capabilities to an autonomous vehicle for planetary exploration. The authors describe the task analysis process that led to the selection of a configuration with three soft fingers. They carry out a complete analysis of the stability of a grasp for this gripper including an analysis of the deformation of the fingers at the points of contact. The implementation of a grasp selection algorithm is described, and results on three-dimensional representations of objects computed from range data are presented.>
C. Francois, Katsushi Ikeuchi, Martial Hebert
ICRA2
1991 Recovering shape in the presence of interreflections
abstract
An algorithm for recovering the shape and reflectance of Lambertian surfaces in the presence of interreflections is presented. The surfaces may be of arbitrary but continuous shape, and with possibly varying and unknown reflectance. The actual shape and reflectance are recovered from the pseudoshape and pseudoreflectance estimated by a local shape-from-intensity method (e.g., photometric stereo). Thus, the algorithm enhances the performance and the utility of existing shape-from-intensity methods. From the results reported, two observations can be made that are pertinent to machine vision: interreflections can cause vision algorithms to produce unacceptably erroneous results and hence should not be ignored; and at least some interreflection problems are tractable and solvable.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICRA2
1991 Trajectory generation with curvature constraint based on energy minimization
abstract
The trajectory generation problem for mobile robots consists in providing a set of trajectories that are 'smooth' and meet certain boundary conditions. The authors present a method to generate curvature continuous trajectories for which the curvature profile is a polynomial function of arc length. An algorithm based on the deformation of a curve by energy minimization allows one to solve general geometric constraints which was not possible by previous methods. Furthermore, it is able to take into account the limitation of radius of curvature of the robot by controlling the extrema of curvature along the path.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
IROS3
1991 Determining linear shape change: Toward automatic generation of object recognition programs,
Katsushi Ikeuchi, Ki-Sang Hong
CVGIP Image Underst.1
1991 Shape from interreflections
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
Int. J. Comput. Vis.2
1991 Determining Reflectance Properties of an Object Using Range and Brightness Images
abstract
The authors discuss a method of recovering reflectance properties of a surface from a range image given by a range finder and a brightness image given by a standard TV camera. The Torrance-Sparrow model is used for the reflectance model. The model consists of the Lambertian and specular components: its reflectance properties consist of the relative strength between the Lambertian and specular components and specular sharpness as well as light source direction. An iterative least square fitting method is used to obtain these parameters based on the range and brightness images. An input image is segmented into four different parts using the parameters: Lambertian reflection, specular reflection, interreflection, and shadow part. The authors also reconstruct ideal images that consist of only Lambertian or specular reflection.>
Katsushi Ikeuchi, Kosuke Sato
IEEE Trans. Pattern Anal. Mach. Intell.1
1991 Introduction to the Special Issue on Physical Modeling in Computer Vision
Takeo Kanade, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 Surface Reflection: Physical and Geometrical Perspectives
abstract
Reflectance models based on physical optics and geometrical optics are studied. Specifically, the authors consider the Beckmann-Spizzichino (physical optics) model and the Torrance-Sparrow (geometrical optics) model. These two models were chosen because they have been reported to fit experimental data well. Each model is described in detail, and the conditions that determine the validity of the model are clearly stated. By studying reflectance curves predicted by the two models, the authors propose a reflectance framework comprising three components: the diffuse lobe, the specular lobe, and the specular spike. The effects of surface roughness on the three primary components are analyzed in detail.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 Modeling sensor detectability with the VANTAGE geometric/sensor modeler
abstract
A G-source, an abstract sensor component representing either a light source or a TV camera, is defined. G-source illumination conditions are investigated to describe the condition under which each G-source illuminates (or observes) surface regions with respect to its illumination/observation directions. Sensor detectability is defined as a function that takes object boundaries and produces a set of points on the object from which the sensor can obtain useful data. It is shown that sensor detectability can be represented by G-source illumination conditions and/or operations between them. The way in which G-sources are combined is described by a sensor-composition tree in which leaf nodes represent G-source illumination conditions and branch nodes represent set operations. Applications of the sensor composition tree to object boundaries are investigated and a set of detectable points is obtained using the VANTAGE geometric modeler. Several 2D structures are proposed to represent these detected points given by VANTAGE. A brief overview of how to use these capabilities in building model-based vision systems is given.>
Katsushi Ikeuchi, Jean-Christophe Robert
IEEE Trans. Robotics Autom.1
1990 Determining reflectance parameters using range and brightness images
abstract
A method is presented for recovering reflectance parameters of optically rough surfaces from a range and a brightness image, both of which are generated by a range-finder. The reflectance model of an optically rough surface consists of two components, known as Lambertian and specular components and respectively represented as a cosine and a Gaussian function, and contains the following three basic parameters: the Lambertian strength, specular strength, and specular sharpness. An iterative least-squares fitting method is used to obtain these parameters derived from range and brightness images. Assuming all pixels only contain the Lambertian component, the authors fit the Lambertian function to all the data points. Based on the fitting of these results, they use a threshold derived from a sensor model of the range-finder and exclude those pixels lying outside the bounds of this threshold. To examine the convergence of the algorithm, the authors implemented this algorithm and applied it to several synthesized images. Several experiments using real images demonstrated the applicability of the algorithm.>
Katsushi Ikeuchi, Kosuke Sato
ICCV1
1990 Shape from interreflections
abstract
An iterative algorithm is presented that simultaneously recovers the actual shape and the actual reflectance from the pseudo estimates. The recovery algorithm works on Lambertian surfaces of arbitrary shape with possibly varying and unknown reflectance. The general behavior of the algorithm and its convergence properties are discussed. Both simulation and experimental results are included to demonstrate the accuracy and stability of the algorithm.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICCV2
1990 Minimum cost aspect classification: a module of a vision algorithm compiler
abstract
The authors present the design of a vision algorithm compiler module for object localization that is used to construct efficient vision programs for a subtask of object localization called aspect classification. Intuitively, an aspect is a representative appearance, and can be associated with a range of viewing positions. In object localization, an object is first classified into an aspect in order to obtain a rough estimate of an object's configuration; this is followed by a numerical minimization procedure to locate the object precisely. The compiler module generates an optimal strategy for aspect classification, in the sense that the average cost of classification is minimal. The performance of the module is illustrated with several examples.>
Ki Sang Hong, Katsushi Ikeuchi, Keith D. Gremban
ICPR (1)2
1990 Determining shape and reflectance of hybrid surfaces by photometric sampling
abstract
A method is presented for determining the shapes of hybrid surfaces without prior knowledge of the relative strengths of the Lambertian and specular components of reflection. The object surface is illuminated using extended light sources and is viewed from a single direction. Surface illumination using extended sources makes it possible to ensure the detection of both Lambertian and specular reflections. Uniformly distributed source directions are used to obtain an image sequence of the object. This method of obtaining photometric measurements is called photometric sampling. An extraction algorithm uses the set of image intensity values measured at each surface point to compute orientation as well as relative strengths of the Lambertian and specular reflection components. The simultaneous recovery of shape and reflectance parameters enables the method to adapt to variations in reflectance properties from one scene point to another. Experiments were conducted on Lambertian surfaces, specular surfaces, and hybrid surfaces.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
IEEE Trans. Robotics Autom.2
1989 Determining linear shape change: toward automatic generation of object recognition programs
abstract
A 3-D object-localization task may be divided into two parts. First, one visible region is classified into one of the aspects of the 3-D object where an aspect is defined as a topologically equivalent class of appearances. Then, the precise attitude and position of the object are determined within one aspect. The authors generate a program to determine the precise attitude and position of an object within one aspect, provided that the face correspondences are given as the result of aspect classification. They establish rules (to define each free coordinate system at each aspect) and correspondences between model edges and image edges, and iteratively solve the transformation equation to determine the object's attitude and position using these correspondences. To turn the strategy into a runnable program, an object library and a geometric compiler are prepared.>
Katsushi Ikeuchi, Ki Sang Hong
CVPR1
1989 Shape and reflectance from an image sequence generated using extended sources
abstract
The authors present a method for determining the shapes of surfaces whose reflectance properties may vary from Lambertian to specular, without prior knowledge of the relative strengths of the Lambertian and specular components of reflection. The object surface is illuminated using extended light sources and is viewed from a single direction. Surface illumination using extended sources makes it possible to ensure the detection of both Lambertian and specular reflections. Multiple source directions are used to obtain an image sequence of the object. An extraction algorithm uses the set of image intensity values measured at each surface point to compute orientation as well as relative strengths of the Lambertian and specular reflection components. The proposed method is called photometric sampling, as it uses samples of photometric function that relates image intensity to surface orientation, reflectance, and light source characteristics. Experiments were conducted on Lambertian surfaces, specular surfaces, and hybrid surfaces, whose reflectance models are composed of both Lambertian and specular components. The results show high accuracy in measured orientations and estimated reflectance parameters.>
Shree K. Nayar, Katsushi Ikeuchi, Takeo Kanade
ICRA2
1989 Modelling sensors: Toward automatic generation of object recognition program
Katsushi Ikeuchi, Takeo Kanade
Comput. Vis. Graph. Image Process.1
1988 Applying Sensor Models To Automatic Generation Of Object Recognition Programs
abstract
One of the most important and systematic methods to build modelbased vision systems is that to generate object recognition programs automatically from given geometric models. Automatic generation of object recognition programs requires several key components to be developed: object models to describe the geometric and photometric properties of an object to be recognized, sensor models to predict object appearances from the object model under a given sensor, strategy generation using the pred,icted appearances to produce an recognition strategy, and program generation converting the recognition strategy into an executable code. This paper concentrates on sensor modeling and its relationship with strategy generation, because we regard it as the bottle neck to automatic generation of object recognition programs. We consider two aspects of sensor characteristics: sensor detectability and sensor reliability. Sensor detectability specifies what kinds of features can be detected and in what condition the features are detected; sensor reliability is a confidence for the detected features. We define the configuration space to represent sensor characteristics. We propose a representation method for sensor detectability and rcliability in the configuration space. Finally, we investigate how to use the proposed sensor modcl in automatic generation of an objcct recognition program.
Katsushi Ikeuchi, Takeo Kanade
ICCV1
1988 Automatic generation of object recognition programs
abstract
Issues and techniques are discussed to automatically compile object and sensor models into a visual recognition strategy for recognizing and locating an object in three-dimensional space from visual data. Automatic generation of recognition programs by compilation, in an attempt to automate this process, is described. An object model describes geometric and photometric properties of an object to be recognized. A sensor model specifies the sensor characteristics in predicting object appearances and variations of feature values. It is emphasized that the sensors, as well as objects, must be explicitly modeled to achieve the goal of automatic generation of reliable and efficient recognition programs. Actual creation of interpretation trees for two objects and their execution for recognition from a bin of parts are demonstrated.>
Katsushi Ikeuchi, Takeo Kanade
Proc. IEEE1
1987 Generating an interpretation tree from a CAD model for 3D-object recognition in bin-picking tasks
Katsushi Ikeuchi
Int. J. Comput. Vis.1
1984 Shape from Regular Patterns
Katsushi Ikeuchi
Artif. Intell.1
1983 Application of 3-D models to computer vision
Yoshiaki Shirai, Kazutada Koshikawa, Masaki Oshima, Katsushi Ikeuchi
Comput. Graph.4
1982 A Model Based Vision System for Recognition of Machine Parts
Katsushi Ikeuchi, Yoshiaki Shirai
AAAI1
1981 Recognition of 3-D Objects Using the Extended Gaussian Image
Katsushi Ikeuchi
IJCAI1
1981 Numerical Shape from Shading and Occluding Boundaries
Katsushi Ikeuchi, Berthold K. P. Horn
Artif. Intell.1
1981 Determining Surface Orientations of Specular Surfaces by Using the Photometric Stereo Method
abstract
The orientation of patches on the surface of an object can be determined from multiple images taken with different illumination, but from the same viewing position. The method, referred to as photometric stereo, can be implemented using table lookup based on numerical inversion of reflectance maps. Here we concentrate on objects with specularly reflecting surfaces, since these are of importance in industrial applications. Previous methods, intended for diffusely reflecting surfaces, employed point source illumination, which is quite unsuitable in this case. Instead, we use a distributed light source obtained by uneven illumination of a diffusely reflecting planar surface. Experimental results are shown to verify analytic expressions obtained for a method employing three light source distributions.
Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.1
1979 An Application of the Photometric Stereo Method
Katsushi Ikeuchi, Berthold K. P. Horn
IJCAI1