Takashi Komuro

dblp:88/3350 · DBLP profile ↗
← Back
48ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 18 · 3 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 1 since 2021Systems, architecture and hardware · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Reproducing the Appearance of Metallic Materials by Capturing Long-Range Dependencies Using Image-to-Image Translation Network with Vision Transformer
Kaito Kojima, Taishi Iriyama, Takashi Komuro
CGI (3)3
2025 AR Digital Workspace: Extending the Workspace of a Mobile Device into Real Space
abstract
We propose AR Digital Workspace that extends the workspace of a mobile device in order to solve the narrow display area of mobile devices. The proposed interface allows a user to place multiple application windows on a flat surface in real space through a mobile device. The display area can be changed by moving the mobile device, allowing the user to easily switch between the applications. The user can change the positions and sizes of windows, operate applications in the windows, and copy and paste text between windows by touching the screen of the mobile device. The user can also copy characters in real space by recognizing characters in the images captured by the device’s camera. This interface allows the user to work with a large amount of information displayed in a large workspace. We conducted an experiment to compare the proposed interface with a traditional touchscreen interface using a task that involves switching between multiple applications to explore necessary information and answer questions (N=14). The results showed no difference in task completion time between the methods, but as trials were repeated, the completion time with the proposed method decreased, and on the last trial, the proposed method took significantly less time than the touchscreen interface. It was also shown that the proposed method was significantly less burdensome for more than half of the NASA-TLX items.
Yuki Kojima, Taishi Iriyama, Takashi Komuro
Int. J. Hum. Comput. Interact.3
2024 Self-measurement of 3D Leg Shape Using a Smartphone Through a Mirror
Yee Win Shwe, Takashi Komuro, Keiko Ogawa-Ochiai, Norimichi Tsumura
ICIC (1)2
2023 View Interpolation Networks for Reproducing Material Appearance of Specular Objects
abstract
In this study, we propose view interpolation networks to reproduce changes in the brightness of an object's surface depending on the viewing direction, which is important in reproducing the material appearance of a real object. We use an original and a modified version of U-Net for image transformation. The networks were trained to generate images from intermediate viewpoints of four cameras placed at the corners of a square. We conducted an experiment with three different combinations of methods and training data formats. We found that it is best to input the coordinates of the viewpoints together with the four camera images and to use images from random viewpoints as the training data.
Chihiro Hoshizawa, Takashi Komuro
Virtual Real. Intell. Hardw.2
2022 Measuring Shape and Reflectance of Real Objects Using a Handheld Camera
Yee Win Shwe, Zar Zar Tun, Seiji Tsunezaki, Takashi Komuro
ICIC (1)4
2022 Blockwise Feature-Based Registration of Deformable Medical Images
Su Wai Tun, Takashi Komuro, Hajime Nagahara
ICIC (1)2
2022 An asymmetrical-structure auto-encoder for unsupervised representation learning of skeleton sequences
Takashi Komuro
Comput. Vis. Image Underst.2
2022 Users' Content Memorization in Multi-User Interactive Public Displays
abstract
In this study, we investigate users’ content memorization in interactive public displays that can be simultaneously used by multiple persons. We designed a system in which interaction would enhance users’ content memorization, and that allows multiple users to collaborate on interaction. We installed the system in a university campus and conducted a field study on passersby. We interviewed the users who continued interacting with the system to reach or pass a certain phase, and evaluated users’ content memorization from the accuracy of displayed information that was unconsciously memorized. As a result, the users who interacted with the display had a significantly higher correct answer rate for detailed questions about the displayed information than those who only watched a video on the display, which suggests that our interactive public display resulted in better content memorization. It is also shown that displayed information was efficiently delivered to many users by enabling multiple persons to interact with the display.
Narumi Sugiura, Rikako Ogura, Yoshio Matsuda, Takashi Komuro, Kayo Ogawa
Int. J. Hum. Comput. Interact.4
2021 Transmission of correct gaze direction in video conferencing using screen-embedded cameras
Kazuki Kobayashi, Takashi Komuro, Keiichiro Kagawa, Shoji Kawahito
Multim. Tools Appl.2
2021 Correction to: Transmission of correct gaze direction in video conferencing using screen-embedded cameras
Kazuki Kobayashi, Takashi Komuro, Keiichiro Kagawa, Shoji Kawahito
Multim. Tools Appl.2
2021 AR Peephole Interface: Extending the workspace of a mobile device using real-space information
Masashi Miyazaki, Takashi Komuro
Pervasive Mob. Comput.2
2020 Recognizing Gestures from Videos using a Network with Two-branch Structure and Additional Motion Cues
abstract
In this paper, we propose a method for recognizing gestures from videos which implicitly incorporates multimodal data during training, and makes classification by only using RGB modality data. The network is 3d-convolutional, and includes a shared network for implicitly incorporating multiple modalities, a generation branch for estimating motion regions and a classification branch for classifying gestures. We introduce a type of efficient modality data, binarized motion cues, which include information of moving hand regions, and are learned by using the generation network. The binarized motion cues are given as extra supervision for learning motion in the generation branch. Since features of additional motion cues learned by the generation branch are implicitly fused with features learned by the classification branch, the classification performance can be improved. Experimental results showed that the shared network can extract more discriminable intermediate features, and the network with the classification branch can achieve improved performance by only using RGB modality input data.
Takashi Komuro
FG2
2020 Dynamic layout optimization for multi-user interaction with a large display
abstract
In this paper, we propose a user interface that allows multiple users to obtain information interactively and to operate the interface simultaneously on a large display. Users can interact with the display using hand gestures, which enables many users to easily use the system from a distance. In order for each user to effectively use the screen and also not to be interfered with by other users, the system dynamically optimizes the layout of the screen. We created a prototype system and conducted an experiment to compare the proposed interface with those without optimization. The results showed that the participants were able to obtain more information per unit time using the proposed interface. Moreover, the participants were less hurried or frustrated, and felt less interfered with by other users.
Yoshio Matsuda, Takashi Komuro
IUI2
2019 Bivariate BRDF Estimation Based on Compressed Sensing
Haru Otani, Takashi Komuro, Shoji Yamamoto, Norimichi Tsumura
CGI2
2019 Semi-Automatic Creation of an Anime-Like 3D Face Model from a Single Illustration
abstract
In this paper, we propose a method for semi-automatically creating an anime-like 3D face model from a single illustration. In the proposed method, principal component analysis (PCA) is applied to existing anime-like 3D models to obtain base models for generating natural 3D models. To align the dimensions of the data and make geometric correspondence, a template model is deformed using a modified Nonrigid Iterative Closest Point (NICP) method. Then, the coefficients of the linear combination of the base models are estimated by minimizing the difference between the rendered image of the 3D model with the coefficients and the input illustration using edge-based matching. We confirmed that our method was able to generate a natural anime-like 3D face models which has similar eye and face shapes to those of the input illustration.
Takayuki Niki, Takashi Komuro
CW2
2019 Recognizing Fall Actions from Videos Using Reconstruction Error of Variational Autoencoder
abstract
In this paper, we propose a method for detecting fall actions using a variational auto-encoder (VAE) with 3D-convolutional residual blocks. The VAE learns a distribution of Activity of Daily Life (ADL) data, and the reconstruction error is used to detect fall actions. The proposed method is a kind of unsupervised learning method with weakly labeled data for solving the problem of imbalance between the amount of fall data and that of ADL data. Furthermore, we propose extracting a human region from an entire image using skeleton information and aligning motions using the same joint point so that the neural network can focus on learning human motions, which can enhance the accuracy of fall detection. The results of experiments showed that our method achieved a competitive level of accuracy and better generalization ability compared to supervised learning with well-labeled data.
Takashi Komuro
ICIP2
2019 An Evaluation of Head-Mounted Virtual Reality for Special Education from the Teachers' Perspective
abstract
In this research, we explore the use of head-mounted virtual reality for special education from the teachers’ perspective. We asked a group of special educators to assess the use of VR headset while students with mental disabilities played a VR game. The teachers concluded that head-mounted VR can be used for teaching students to follow instruction and training for work.
Sirisilp Kongsilp, Takashi Komuro
VRST2
2018 On-mouse projector: Peephole interaction using a mouse with a mobile projector
Tomohiro Araki, Takashi Komuro
Pervasive Mob. Comput.2
2017 Comparative study on text entry methods for mobile devices with a hover function
abstract
In this paper, we compared text entry methods for mobile devices such as smartphones, in which part of a software keyboard around the finger position is enlarged by using a hover function. We examined four methods: fixing the center of enlargement, not fixing the center of enlargement, adding a scrolling function, and using a standard software keyboard for reference. We recruited 12 participants and evaluated the text entry speed and error rates. The method in which the center of enlargement was fixed had the slowest text entry speed, but showed an error rate that was significantly lower than the method in which the center of enlargement was not fixed. This result shows that fixing the center of enlargement enables users to select keys stably and enter target letters accurately.
Toshiaki Aiyoshizawa, Takashi Komuro
MUM2
2017 On-mouse projector: peephole interaction using a mouse with a projector
abstract
In this paper, we propose a system that combines a mouse with a projector and that allows users to perform peephole interaction, which enables stable operation in a wide workspace. In our prototype system, a mobile projector is attached above a mouse and projects part of the workspace on the plane in front of the mouse. The user can change the area he/she wishes to see by moving the mouse and can perform operations in a virtually extended workspace. We conducted an experiment to compare the methods of changing the displayed region and the cursor position when the mouse moves. Our results show that the method with the highest ratio of view transition to mouse movement, in which the workspace is not fixed to the real space, outperformed the other methods and was also highly evaluated by participants. Furthermore, the task completion time decreased with repeating trials with fixed target positions, which suggests that spatial memory is helpful even when the workspace is not fixed to the real space.
Tomohiro Araki, Takashi Komuro
MUM2
2017 Mobile augmented reality for providing perception of materials
Ryota Nomura, Yuko Unuma, Takashi Komuro, Shoji Yamamoto, Norimichi Tsumura
MUM3
2017 Recognition of typing motions on AR typing interface
abstract
We have proposed Augmented Reality (AR) Typing Interface as a new input interface for mobile devices. By using this interface, users can utilize a wide space in the air in inputting texts. This interface uses a camera which is attached to the back of the device. A virtual keyboard is overlaid on the image captured by a camera using AR technology. In this paper, we propose a new method to recognize typing motions more precisely. In this method, feature vectors are generated by using information of time-series optical flows. Typing motions are recognized by using support vector machine (SVM). We confirmed that the proposed method was able to recognize typing motions with about 90% accuracy for the recorded video of typing in the air.
Masae Okada, Masakazu Higuchi, Takashi Komuro, Kayo Ogawa
MUM3
2017 Distant Pointing User Interfaces based on 3D Hand Pointing Recognition
abstract
In this paper, we propose a system that realizes remote control of a computer with small hand gestures. Distant pointing is realized using a 3D hand pointing recognition algorithm that obtains position and direction of the pointing hand. We show the effectiveness of the system by constructing three types of user interfaces considering the accuracy of distant pointing in the current system. We created a tile layout interface for rough selection operation, a pie menu interface for detailed operation, and a viewer interface for document browsing.
Yutaka Endo, Dai Fujita, Takashi Komuro
ISS3
2017 A gaze-preserving group video conference system using screen-embedded cameras
Kazuki Kobayashi, Takashi Komuro, Keiichiro Kagawa, Shoji Kawahito
VRST2
2016 Space-sharing AR interaction on multiple mobile devices with a depth camera
abstract
In this paper, we propose a markerless augmented reality (AR) system that works on multiple mobile devices. The relative positions and orientations of the devices and their individual motions are estimated from 3D information in real space obtained by depth cameras attached to the devices. The system allows multiple users to share the AR space and to interact with the same virtual object. To estimate the relative positions and orientations of the devices, the system generates 2D images by looking down from above at the 3D scene obtained by the depth cameras, performs 2D registration using template matching, and obtains a transformation matrix that transforms the coordinate system of one camera to that of another camera. The motion of a camera is estimated using the ICP algorithm to realize markerless AR. Using the proposed system, we created an application that enables multiple users to interact with the same virtual object.
Yuki Kaneto, Takashi Komuro
VR2
2015 Overlaying Navigation Signs on a Road Surface Using a Head-Up Display
abstract
In this paper, we propose a method for overlaying navigation signs on a road surface and displaying them on a head-up display (HUD). Accurate overlaying is realized by measuring 3D data of the surface in real time using a depth camera. In addition, the effect of head movement is reduced by performing face tracking with a camera that is placed in front of the HUD, and by performing distortion correction of projection images according to the driver's viewpoint position. Using an experimental system, we conducted an experiment to display a navigation sign and confirmed that the sign is overlaid on a surface. We also confirmed that the sign looks to be fixed on the surface in real space.
Kaho Ueno, Takashi Komuro
ISMAR2
2015 Natural 3D Interaction Using a See-Through Mobile ASystem
abstract
In this paper, we propose an interaction system in which the appearance of the image displayed on a mobile display is consistent with that of the real space and that enables a user to interact with virtual objects overlaid on the image using the user's hand. The three-dimensional scene obtained by a depth camera is projected according to the user's viewpoint position obtained by face tracking, and the see-through image whose appearance is consistent with that outside the mobile display is generated. Interaction with virtual objects is realized by using the depth information obtained by the depth camera. To move virtual objects as if they were in real space, virtual objects are rendered in the world coordinate system that is fixed to a real scene even if the mobile display moves, and the direction of gravitational force added to virtual objects is made consistent with that of the world coordinate system. The former is realized by using the ICP (Iterative Closest Point) algorithm and the latter is realized by using the information obtained by an accelerometer. Thus, natural interaction with virtual objects using the user's hand is realized.
Yuko Unuma, Takashi Komuro
ISMAR2
2015 Dynamic 3D interaction using an optical See-through HMD
abstract
We propose a system that enables dynamic 3D interaction with real and virtual objects using an optical see-through head-mounted display and an RGB-D camera. The virtual objects move according to physical laws. The system uses a physics engine for calculation of the motion of virtual objects and collision detection. In addition, the system performs collision detection between virtual objects and real objects in the three-dimensional scene obtained from the camera which is dynamically updated. A user wears the device and interacts with virtual objects in a seated position. The system gives users a great sense of reality through an interaction with virtual objects.
Nozomi Sugiura, Takashi Komuro
VR2
2015 Three-dimensional VR interaction using the movement of a mobile display
abstract
In this study, we propose a VR system for allowing various types of interaction with virtual objects using an autostereoscopic mobile display and an accelerometer. The system obtains the orientation and motion information from the accelerometer attached to the mobile display and reflects them to the motion of virtual objects. It can present 3D images with motion parallax by estimating the position of the user's viewpoint and by displaying properly projected images. Furthermore, our method enables to connect the real space and the virtual space seamlessly through the mobile display by determining the coordinate system so that one of the horizontal surfaces in the virtual space coincides with the display surface. To show the effectiveness of this concept, we implemented an application to simulate food cooking by regarding the mobile display as a frying pan.
Takashi Komuro
VR2
2013 AR typing interface for mobile devices
abstract
We propose a new user interface system for mobile devices. By using augmented reality (AR) technology, the system overlays virtual objects on real images captured by a camera attached to the back of a mobile device, and the user can operate the mobile device by manipulating the virtual objects with his/her hand in the space behind the mobile device. This system allows the user to operate the device in a wide three-dimensional space and to select small objects easily. Also, the AR technology provides the user with a sense of reality in operating the device. We developed a typing application using our system and verified the effectiveness by user studies. The results showed that more than half of the subjects felt that the operation area of the proposed system is larger than that of a smartphone and that both AR and unfixed key-plane are effective for improving typing speed.
Masakazu Higuchi, Takashi Komuro
MUM2
2011 Stereo 3D reconstruction using prior knowledge of indoor scenes
abstract
We propose a new method of indoor-scene stereo vision that uses probabilistic prior knowledge of indoor scenes in order to exploit the global structure of artificial objects. In our method, we assume three properties of the global structure - planarity, connectivity, and parallelism/orthogonality - and we formulate them in the framework of maximum a posteriori (MAP) estimation. To enable robust estimation, we employ a probability distribution that has both high peaks and wide flat tails. In experiments, we demonstrated that our approach can estimate shapes whose surfaces are not constrained by three orthogonal planes. Furthermore, comparing our results with those of a conventional method that assumes a locally smooth disparity map suggested that the proposed method can estimate more globally consistent shapes.
Kentaro Kofuji, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICRA3
2011 VolVision: high-speed capture in unconstrained camera motion
abstract
In this paper, we propose a novel concept called VolVision that encompasses using a camera to reconstruct 6DoF unconstrained motion. "VolVision" is designed to handle imagery falling, tossed or thrown cameras. And VolVision also allows users to reconstruct dynamic images and generate a 3D-mapped scene from image sequences. It could be used to model severe environments like valley and mountains that are normally not easily viewed by humans. We produced a prototype that embodies the concept above, and were able to reconstruct the camera's path, perform image mosaicing, and track 3D information of feature points in images.
Hideki Takeoka, Yushi Moko, Carson Reynolds, Takashi Komuro, Yoshihiro Watanabe, Masatoshi Ishikawa
SIGGRAPH Asia Sketches4
2011 Human gait estimation using a wearable camera
abstract
We focus on the growing need for a technology that can achieve motion capture in outdoor environments. The conventional approaches have relied mainly on fixed installed cameras. With this approach, however, it is difficult to capture motion in everyday surroundings. This paper describes a new method for motion estimation using a single wearable camera. We focused on walking motion. The key point is how the system can estimate the original walking state using limited information from a wearable sensor. This paper describes three aspects: the configuration of the sensing system, gait representation, and the gait estimation method.
Yoshihiro Watanabe, Tetsuo Hatanaka, Takashi Komuro, Masatoshi Ishikawa
WACV3
2010 Wide range image sensing using a thrown-up camera
abstract
In this paper, we propose a wide-range image sensing method using a camera thrown up into the air. By using camera thrown up in this way, we can get images that are otherwise difficult to obtain, such as those taken from overhead. As an example of wide-range image sensing, we integrated video images captured by a thrown-up camera using an image mosaicing technique. When rotation about the optical axis of the camera can be ignored, we can integrate images by mosaicing using a translational approximation, which preferentially pastes pixels around the image center. To obtain the information about the camera direction, a rotational approximation using the angles of incident light rays is required. We also propose use of a high-frame-rate camera (HFR camera) in order to acquire a large amount of information. A seamless large image was obtained by synthesizing the images captured by a thrown-up HFR camera. We found that high frame rates of around 1000 fps were necessary.
Toshitaka Kuwa, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICME3
2010 Estimation of Non-rigid Surface Deformation Using Developable Surface Model
abstract
There is a strong demand for a method of acquiring a non-rigid shape under deformation with high accuracy and high resolution. However, this is difficult to achieve because of performance limitations in measurement hardware. In this paper, we propose a model based method for estimating non-rigid deformation of a developable surface. The model is based on geometric characteristics of the surface, which are important in various applications. This method improves the accuracy of surface estimation and planar development from a low-resolution point cloud. Experiments using curved documents showed the effectiveness of the proposed method.
Yoshihiro Watanabe, Takashi Nakashima, Takashi Komuro, Masatoshi Ishikawa
ICPR3
2010 Surface image synthesis of moving spinning cans using a 1, 000-fps area scan camera
Tomohira Tabata, Takashi Komuro, Masatoshi Ishikawa
Mach. Vis. Appl.2
2010 A Reconfigurable Embedded System for 1000 f/s Real-Time Vision
abstract
In this paper, we proposed an architecture of embedded systems for high-frame-rate real-time vision on the order of 1000 f/s, which achieved both hardware reconfigurability and easy algorithm implementation while fulfilling performance demands. The proposed system consisted of an embedded microprocessor and field programmable gate arrays (FPGAs). A coprocessor consisting of memory units, direct memory access controller units, and image processing units were implemented in each FPGA. While the number of units and functions are reconfigurable by reprogramming the FPGAs, users can implement algorithms without hardware knowledge. A descriptor method in which the central processing unit gave instructions to each coprocessor through a register array enabled task-level parallel processing as well as pixel-level parallel processing in the processing units. The specifications of an evaluation system developed based on the proposed architecture, the results of performance evaluation, and application examples using the system were shown.
Takashi Komuro, Tomohira Tabata, Masatoshi Ishikawa
IEEE Trans. Circuits Syst. Video Technol.1
2009 High-resolution shape reconstruction from multiple range images based on simultaneous estimation of surface and motion
abstract
Recognition of dynamic scenes based on shape information could be useful for various applications. In this study, we aimed at improving the resolution of three-dimensional (3D) data obtained from moving targets. We present a simple clean and robust method that jointly estimates motion parameters and a high-resolution 3D shape. Experimental results are provided to illustrate the performance of the proposed algorithm.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICCV2
2008 High-S/N imaging of a moving object using a high-frame-rate camera
abstract
In this paper we propose a high-S/N imaging method involving combining many images captured with small blur using a video camera capable of high-frame-rate image capturing at 1000 frames/s. Use of a high-frame-rate camera makes the image change between frames small, enabling easy motion estimation, and makes it possible to use more light information, even when the exposure time is reduced to avoid blurring. To obtain a clear picture without misalignment due to motion parallax, it is necessary to determine both the motion and a depth map of the subject from noisy input images. We show results when applying the proposed algorithm to an image sequence captured by a high-frame-rate camera.
Takashi Komuro, Yoshihiro Watanabe, Masatoshi Ishikawa, Tadakuni Narabu
ICIP1
2008 Integration of time-sequential range images for reconstruction of a high-resolution 3D shape
abstract
The recognition of dynamic scenes using 3D shapes could provide useful approaches for various applications. However, the conventional 3D-shape sensing systems dedicated for such scenes have had problems in spatial resolution, though they have achieved high sampling rate in temporal domain. In order to solve this limits, we present a method that integrates time-sequential partial range images capturing moving targets to reconstruct a high-resolution range image. In the proposed method, multiple range images are set in the same coordinate system based on multi-frame simultaneous alignment. This paper also demonstrates the performance of the proposed method using some example rigid bodies.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICPR2
2007 A High-Speed Vision System for Moment-Based Analysis of Numerous Objects
abstract
We describe a high-speed vision system for real-time applications, which is capable of processing visual information at a frame rate of 1 kfps, including both imaging and processing. Our system performs moment-based analysis of numerous objects. Moments are useful values providing information about geometric features and invariant features with respect to image-plane transformations. In addition, the simultaneous observation of numerous objects allows recognition of various complex phenomena. The proposed system achieves high-speed image processing by providing a dedicated massively parallel co-processor for moment extraction. The co-processor has a high-performance core based on a pixel-parallel and object-parallel calculation method. We constructed a prototype system and evaluated its performance. We present results obtained in actual operation.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICIP (5)2
2007 A Moment-based 3D Object Tracking Algorithm for High-speed Vision
abstract
In this paper we propose a method of realizing continuous tracking of a three-dimensional object by calculating moments of a translating and rotating object whose shape is known, either analytically or by using a table, and matching them with those of the input image. In simulation, the position and orientation is accurately recognized. Using a noise model and particle filter, we show that the position and orientation can be recognized even with noisy images.
Takashi Komuro, Masatoshi Ishikawa
ICRA1
2007 955-fps Real-time Shape Measurement of a Moving/Deforming Object using High-speed Vision for Numerous-point Analysis
abstract
This paper describes real-time shape measurement using a newly developed high-speed vision system. Our proposed measurement system can observe a moving/deforming object at high frame rate and can acquire data in real-time. This is realized by using two-dimensional pattern projection and a high-speed vision system with a massively parallel co-processor for numerous-point analysis. We detail our proposed shape measurement system and present some results of evaluation experiments. The experimental results show the advantages of our system compared with conventional approaches.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICRA2
2007 Design of a Massively Parallel Vision Processor based on Multi-SIMD Architecture
abstract
Increasing demands for robust image recognition systems require vision processors not only with enormous computational capacities but also with sufficient flexibility to handle highly complicated recognition tasks. We describe a multi-SIMD architecture and the design of a vision processor based on it for carrying out such difficult image recognition tasks. The proposed architecture consists of two SIMD parallel processing modules and a shared memory, allowing highly parallelized and flexible computation of complicated recognition tasks, which were difficult to process on a conventional massively parallel SIMD architecture. We designed a prototype vision processor for evaluation purposes and confirmed that the processor could be implemented in FPGA.
Kota Yamaguchi, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ISCAS3
2005 Safe human-robot-coexistence: emergency-stop using a high-speed vision-chip
abstract
The coexistence of humans and industrial robots in a common workspace provides the advantage of increased flexibility in production or longer system up-time during maintenance. However, it is fundamentally necessary to guarantee the safety of the human. This paper presents an approach that uses a specialized tracking-vision-chip to realize a high-speed emergency-stop for safe human-robot-coexistence. The presented approach uses the ability of the vision-chip to perform pixel-parallel masking and fast summation-operations on binary images to detect whether the robot and human are too close. After initial evaluation in a (semi)-simulation, the approach was realized in an experimental system. Even with only a small 8-bit microcontroller controlling the vision-chip and the communication, a cycle time of more than 500Hz was achieved.
Dirk M. Ebert, Takashi Komuro, Akio Namiki, Masatoshi Ishikawa
IROS2
2002 A Real-Time Visual Processing System using a General-Purpose Vision Chip
abstract
A real-time visual processing system using a general-purpose vision chip, an image sensor in which photo detectors and processing elements are integrated, is described. In order to control the vision chip and process its output at high speed, a novel architecture called SPARSIS, in which the control process of the vision chip is pipelined and integrated with a RISC type integer pipeline, was developed. This architecture can guarantee real-tune operation with high temporal resolution, and even makes possible software-controlled A/D conversion. Sample algorithms demonstrating its fine-grained real timeliness, and experimental results with the implemented system, are also described.
Shingo Kagami, Takashi Komuro, Idaku Ishii, Masatoshi Ishikawa
ICRA2
2002 High-speed sensory-motor fusion based on dynamics matching
abstract
This paper discusses a design concept of a sensory-motor fusion system to achieve high performance in a dynamic changing environment. From the viewpoint of a dynamic system, the new concept called "dynamics matching" is proposed to match the dynamics constraints of a system. Based on this concept, we describe a high-speed vision chip that has a general purpose parallel processing array along with a photodetector all in a single silicon chip. Next we describe a new sensory-motor fusion system which consists of a hierarchical parallel processing system, a vision chip system, and a multifingered hand-arm. All sensory feedback, including visual feedback, can be achieved in 1 ms. In addition, as an application of the system, we demonstrate high-speed grasping using visual and force feedback.
Akio Namiki, Takashi Komuro, Masathoshi Ishikawa
Proc. IEEE2
1992 A distributed cooperative CASE environment for communications software
abstract
The authors propose a development process model and support environment for developing intelligent business communication services and other evolving communication services. To increase customer satisfaction and cut development time, they have developed a development process model, distributed concurrent development, in which multiple services are concurrently developed at a number of regionally distributed development centers. To support the process model, the ICAROS tool was developed, where developers and customers can cooperatively define service specifications with a visual object-oriented specification language. To smoothly migrate existing communication service to assets into intelligent network services, ICAROS supports the cyclic process model, which integrates forward engineering and reverse engineering, together with amplifier concept, which augments human ability on a groupware platform. Experimental use of ICAROS revealed the positive effect of proposed approach.>
Mikio Aoyama, Masami Nakamura, Shinya Kawajiri, Kousuke Takahashi, Takanori Hashizume, Takashi Komuro
COMPSAC6