Sota Shimizu

dblp:14/5586 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
10since 2021 · last 2025
0009-0006-2296-4443ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Approximate Representation of Gaze Rate using GMM
abstract
In this paper, the authors aim at representing a gaze rate distribution approximately using Gaussian Mixture Model (GMM). The gaze rate, i.e., a time-series of distributions of gaze probabilities is calculated statistically from gaze point data measured from many audiences. It is very effective to analyze when and where the audience paid attention to in a target movie. In recent days when TV becomes more personalized as internet streaming service gets broadly-spread. Thus, the gaze rate is expected as a novel strong criterion to evaluate more accurately how much degree TV commercials merchandise to the audience, compared to a well-known viewer rating. However, when its applications are taken into account, its calculation cost of a rigid gaze rate data was too high, and its data amount was too large to transmit and process it in a real time data. Therefore, this paper proposes approximate representation of the gaze rate using GMM in order to solve the above problems of the calculation cost and the data amount. Verification experiments and discussions have been conducted in a comparison between the rigid gaze rate calculated by the existing method and our proposed approximate representation using GMM by paying attention to the calculation cost and accuracy of the approximate gaze rate.
Sota Shimizu, Miwa Takase, Takumi Morimoto
IECON1
2024 Gaze-Based Intention Recognition for Human-Robot Collaboration
abstract
This work aims to tackle the intent recognition problem in Human-Robot Collaborative assembly scenarios. Precisely, we consider an interactive assembly of a wooden stool where the robot fetches the pieces in the correct order and the human builds the parts following the instruction manual. The intent recognition is limited to the idle state estimation and it is needed to ensure a better synchronization between the two agents. We carried out a comparison between two distinct solutions involving wearable sensors and eye tracking integrated into the perception pipeline of a flexible planning architecture based on Hierarchical Task Networks. At runtime, the wearable sensing module exploits the raw measurements from four 9-axis Inertial Measurement Units positioned on the wrists and hands of the user as an input for a Long Short-Term Memory Network. On the other hand, the eye tracking relies on a Head Mounted Display and Unreal Engine.
Valerio Belcamino, Miwa Takase, Mariya Kilina, Alessandro Carfì, Fulvio Mastrogiovanni, Akira Shimada, Sota Shimizu
AVI7
2024 Challenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder
Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai
INTERSPEECH3
2023 Falcon: Wide Angle Fovea Vision System for Marine Rescue Drone
abstract
In this paper, we propose a new vision system for marine rescue drones called Falcon. Our system is designed to track and monitor individuals in need of rescue using a wide-angle field of view and a gimbal mechanism to control the camera direction. Additionally, our system utilizes a high-resolution central field of view to assess the priority level of the person in need of rescue. We conducted two experiments to test our system's performance. The first experiment focused on the correlation between the data compression rate of the system and the precision and recall scores of human detection using YOLO in three different flight altitudes. The second experiment examined the correlation between the distance from the center of the images captured by the system and the confidence score of the detected persons, considering the resolution difference from the central region to the periphery.
Tetsuya Oda, Sota Shimizu, Rikuto Nakamoto, Alessandro Carfì, Fulvio Mastrogiovanni
IECON2
2022 Tornado: Power Assist Suit to Assist Twisting Motion of Lower Back
abstract
In this paper, the authors develop a 2-DOF power assist suit (PAS), namely Tornado. Our developed PAS assists power for twisting motions in addition to lifting and lowering motions in environments of agriculture, fishery, factory, and so on. Tornado PAS not only assists power to some motions but also protects our lower back by limiting high-risk motions which causes lower back injuries. The authors apply a differential gear box mechanism to achieve 2-DOF power assist. The prototype of 2-DOF PAS is implemented by comprising two DC motor, two rotary encoders, two current control motor drivers, 2-channel encoder counter board, 2-channel DIO board, battery, and Raspi4 controller installing Linux OS. Although Tornado PAS assists power according to an external force, our system do not use any force sensor. We apply a disturbance observer (DOB) to estimate such an external force. Thus, Tornado is characterized as follows: (1) a unique mechanism using a differential gear box, (2) not only power assist but also prohibiting dangerous motions by computer control, and (3) force sensorless power, i.e., force, assist. In this paper, the authors model Tornado, and establish equation of motion in order to simulate various types of control methods and observers. We experiment power assist of twisting motions based on some position tracking control. Results this control and values estimated using DOB are discussed and evaluated for effective power assist for twisting motion.
Motoki Hirose, Sota Shimizu, Rikuta Mazaki
IECON2
2022 Sidewinder: Snake Robot's Stereo Vision System for Rescue in Collapsed Debris at Disaster Sites
abstract
In this paper, the authors present Sidewinder, a unique stereo vision system for snake-shaped robots operating in search and rescue scenarios. The robot should navigate the environment by finding and passing through narrow spaces (i.e., forward vision task) and search for persons in need of rescue (i.e., panoramic vision task). Therefore, we propose two types of image mapping methods to generate respectively input images for stereo visual SLAM and person detection. This work presents preliminary results using Sidewinder's images as input for the ORS-SLAM2 for localization and mapping, and YOLO-v3, for person detection.
Rikuto Nakamoto, Sota Shimizu, Tomoki Takamura, Alessandro Carfì, Fulvio Mastrogiovanni
IECON2
2022 Gaze Preference Decision Making Predictor Using RNN Classifier
abstract
This paper presents a gaze preference decision making predictor using a recurrent neural network (RNN) classifier from human eye movement when someone needs to choose one of two targets. When the two target objects are displayed side by side on a screen of a VR head mount display (HMD), this predictor predicts his/her final decision in real time by inputting data of the eye movement to the RNN classifier. It is well known that a chosen likelihood increases rapidly approximately 0.5s in advance from when the participants made decision. This result is based on a statistic analysis. In other words, this is not prediction but postdiction. This phenomenon is called the gaze cascade effect. This effect means that people has a positive feedback loop, i.e., the more someone looks at something he/she likes better by choosing one from two candidates according to left or right eye movement, the better he/she likes it. This paper aims at achieving not to postdict but to predict someone’s decision making in real time. In the experiments, we used Blu-ray packages as a target of the gaze preference decision making. First, two packages were chosen randomly from the 78 Blu-ray ones, and were displayed on the VR HMD side by side. 220 sets of eye movement data were measured from 11 participants by an eye-tracking device equipped on the VR HMD. They were recorded with the participants’ final decision of left or right for training the RNN classifier. The 11 participants were instructed to choose one package they like better. Each participant repeated this trial 20 times. With respect to the experimental results, the authors classified all the eye movement data into three categories according to types of movement and data length. Our RNN classifier predicted the participants’ decision making successfully by approximately 91 percent, when the eye data were longer than 3 seconds and include a switching motion between left and right.
Shumpei Sato, Sota Shimizu, Koh Hamada
IECON2
2022 2-DOF Haptic Feedback Control Stick for Remote Rover Navigation
abstract
This paper presents development of a 2DOF haptic feedback control stick for remote control of a two-parallel-wheeled rover. In bilateral control, when dynamics are quite different between a leader and follower devices, e.g., between a linear-slider type of control stick and a wheel with a comparatively large moment of inertia, a force reproduced on the control stick based on external forces applied to the rover’s wheel, is disturbed largely. Since the position control forces the control stick’s position and the rover’s angular velocity to be synchronized, the operator feels not only the reproduced force but also a force by this position control inevitably on the same time. We call this force a restoring force. The restoring force disturbs the operator to feel the reproduced force accurately. But the authors think the restoring force is meaningful for remote navigation of the rover as another haptic information. Thus, in this paper, we developed a 2-DOF haptic control stick, which transfers the reproduced force and the restoring force separately to a slider and a grip, and proposed a control method for achieving this purpose. The control stick was implemented and was performed successfully in verification experiments using the proposed control method. Experimental results were discussed comparing to the former 1-DOF control stick.
Tomonori Yamazaki, Sota Shimizu, Rikuta Mazaki, Hokuto Kurihara, Naoki Motoi, Roberto Oboe, Nobuyuki Hasebe, Tomoyuki Miyashita
IECON2
2021 Effective Voltage Control of Liquid Crystal Lens for Rapid Focal Length Change
abstract
A liquid crystal (LC) lens can make any lens state from positive (convex) to negative (concave) by applying an external voltage. However, when the LC lens changes its focal length, it is difficult to say its response time is fast enough. In this paper, the authors aim at making the response time of the LC lens be shorter by using a control theory to a value of the external voltage applied to the LC lens. We constructed an automatic control system which can change the focal length faster by a computer program. The computer program gives appropriate effective values of AC voltages with the LC lens based on a feed forward control. We conducted experiments to change the focal length of the LC lens from a non-lens state to 200mm and -100mm, respectively. When our proposed effective voltage control method was performed, the response time was successfully reduced in half, compared to a case when a simple step-like voltage was applied.
Tsugumi Fukui, Sota Shimizu, Keigo Muryobayashi, Marenori Kawamura, Susumu Sato, Nobuyuki Hasebe
IECON2
2021 A Visual Odometry for Wide Angle Fovea Sensor SLAM
abstract
This paper presents a method of wide angle fovea visual odometry (WAF-VO) for Wide Angle Fovea Sensor SLAM (WAF-SLAM), by which a unique locally-high accurate and wide-angle map is generated in addition to camera motion estimation. The WAF sensor is a special-made wide-angle sensor that is inspired from human visual function, i.e., the spatial resolution of the image is not uniform throughout the entire field of view (FOV); it is much higher in the central FOV and decreases rapidly towards the periphery. Our visual odometry method is strongly characterized by a wide-angle FOV and space-variant resolution of the input image from the WAF sensor. A locally-high accurate and wide-angle mapping method is proposed as a major part for WAF-SLAM together with the camera motion estimation. Our proposed method estimates camera motions more stably using very low-spatial-resolution wide-angle images remapped from the input image of the WAF sensor. Using the estimated camera motions, narrow-angle high accurate maps are generated from corresponding feature points in high-spatial resolution central regions of the input image. Wide-angle maps are generated from ones in middle-spatial-resolution wide-angle images remapped from the input image apart from the above very low-spatial-resolution images. When the wide-angle maps are generated, the number of extracted feature points is increased by adjusting contrast threshold values of SIFT feature according to regions of the FOV. A KNN matching method improved using epipolar constraint is proposed and employed for avoidance of mismatching the increased feature points. Thus, the wide-angle maps are generated from more correct corresponding feature points. Finally, the above two types of maps are combined into the unique locally-high accurate and wide-angle map, i.e., a WAF map. Using our proposed method, the WAF map was generated by verification experiments. Furthermore, the paper presents an evaluation of the accuracy and precision of the generated map.
Tomoki Takamura, Sota Shimizu, Rei Murakami, Alessandro Carfì, Fulvio Mastrogiovanni
IECON2
2019 Development of Nonmechanical Zoom Lens System using Liquid Crystal
abstract
The Liquid Crystal (LC) zoom lens system has strong advantages with respect to its size and small electric power consumption, because this lens system does not have any mechanical part. The authors take into consideration a structure of the Galilean telescopic type of the zoom lens system because it makes a total length of the lens system be smaller. Its performance is simulated. There exists a trade-off between the response time of the LC lens and its potential maximum lens power. The lens power, defined as the reciprocal of the focal length, is a very important factor of the LC lens, by which a variable range of the magnification change is determined. In this paper, the authors discussed and tested a structure of the Galilean LC lens zoom system. This structure achieves the same variable range of the magnification change by using the LC lens with a much smaller lens power. This result is quite helpful for the response time and brightness of the LC zoom lens system. Simulation results showed its potentials.
Haruka Hirai, Sota Shimizu, Takumi Saito, Hokuto Kurihara, Marenori Kawamura, Susumu Sato
IECON2
2018 Vision System with High Performance Wide Angle Fovea Lens
abstract
This paper designs and produces a single camera head, i.e., the camera view direction control device, for the high-performance Wide Angle Fovea (WAF) sensor. Since the input image by the WAF sensor has an explicit attention region inside its wide-angle field of view (FOV), the view direction control of the WAF sensor is strongly required to acquire visual data most efficiently from the environments. Further, this paper proposes an algorithm to detect self-motion of the WAF sensor using the optical flow calculated from two temporally-sequential images. We have implemented and experimented this algorithm. The experimental results have proved the proposed algorithm performs well to detect self-motion using only the peripheral FOV. This vision system enables to distinguish the self-motion from detecting moving objects in the still scene. The authors hope this WAF vision system will play a role of eyes and will help to control robots more intelligently in the very near future.
Rei Murakami, Sota Shimizu, Nobuyuki Hasebe
ETFA2
2018 Generation of Multi-Level Disparity Map from Stereo Wide Angle Fovea Vision System
abstract
This paper proposes a novel concept of a multi-level disparity map (distance image) characterized by input images from the Wide Angle Fovea (WAF) lens. The authors focuses on a unique property of the WAF lens, i.e., the lens magnification is the highest locally in the central field of view (FOV) and decreases rapidly towards the peripheral FOV. Thus, the input image by the WAF lens achieves more accurate and wider-angle observation simultaneously without increasing the number of image pixels. Our proposed algorithm generates the multi-level disparity map based on the parallel stereo vision method. By using the disparity map, generated from the central region having the high-spatial resolution, we can measure the distance of a target object being quite far away ahead from the WAF stereo vision system very accurately. On the same time, by using the disparity map generated from the peripheral region, we can obtain 3D information of a wider space comparatively close to the vision system by adequate accuracy to some degree. We have implemented the proposed algorithm to the WAF stereo vision system and have experimented in order to verify its performance. The authors think our proposed multi-level disparity map is quite applicable for improving safety of the automatic driving assistance system.
Naoaki Kameyama, Sota Shimizu, Rei Murakami, Motonori Tominaga, Osamu Shimomura, Yusuke Akamine, Naoki Kawasaki, Kazuhisa Ishimaru, Seiichi Mita
IECON2
2018 Saliency Map for Wide Angle Fovea Vision Sensor
abstract
This paper proposes and develops a saliency map suitable for the wide angle fovea (WAF) sensor. The saliency map is well-known as a computational method inspired from the human cognitive vision processing. This bottom-up processing method is often available for finding a target., which should be paid attention to., automatically from the arbitrary scene. However., it is not necessarily sufficient to use the existing saliency sap directly with the WAF sensor. Usually., we utilize the WAF sensor by combining different-level image processing tasks using its high spatial resolution central field of view (FOV) or its wide-angle FOV cooperatively., because the WAF sensor does not provide with a uniform resolution input image. When we apply the saliency map for the input image by the WAF sensor., we need to take into account unique properties of this biologically-inspired special vision sensor. Therefore., we design a novel saliency map model which is more suitable for the WAF sensor. After configuration of this specific saliency map., some verification experiments were implemented to enhance advantages of our proposed saliency map for the WAF sensor., i.e., WAF saliency map.
Rei Murakami, Sota Shimizu, Tatsuya Yamazaki, Nobuyuki Hasebe
IECON2
2018 Development of Wide Angle Fovea Lens for High-Definition Imager Over 3 Mega Pixels
abstract
This paper presents a high-quality wide-angle fovea lens, i.e., the WAF lens, for the autonomous robot's and vehicle's super-sensing vision system. The WAF lens is well-known in the field of robotic vision with respect to its unique design concept, biologically-inspired from a visual system of the primates. The WAF lens achieves the following two conflicting properties in imaging simultaneously: (1) wide field of view (FOV) and (2) high magnification factor (although only the central FOV achieves it partially). In this paper, the authors designs the WAF lens for the high-resolution photosensitive imaging chip more than 3M pixels. For this design, we decide the following targets on the assumption of applying this WAF lens for the stereo vision system: (1) The WAF lens can measure a very far distance over 100m ahead from the imager accurately. (2) The WAF lens can observe approximately 100-degree wide FOV on the same time. We produce a prototype of this WAF lens with much higher optical performance than our previous developments. The compound system of the prototype includes four aspherical surfaces in its front part to project enough bright images so that the WAF lens is available not only at daytime but also in dark situations at night. The authors experiment and demonstrate the projection tests using the prototype, and discuss about the results as the inspection of this challenging development.
Sota Shimizu, Rei Murakami, Motonori Tominaga, Yusuke Akamine, Naoki Kawasaki, Osamu Shimomura, Kazuhisa Ishimaru, Seiichi Mita
IROS1
2017 Development of wide angle fovea telescope with wide-field-of-view immersive eyepiece
abstract
The authors have developed the wide angle fovea (WAF) telescope. This development was originally motivated from our mind to assist the ranger to save the people floating on the sea. Cruel Tsunami disasters often cause such miserable situations. An objective lens part of this special telescope was formerly developed using the Advanced WAF (AdWAF) model, by which we can design a distribution of the lens curve very flexibly. This special-made objective lens has an enough wide field of view, and thus it can keep a moving target as being always inside field of view. In addition, it can observe the target more in detail because its central field of view has higher magnification than the conventional telescope. Yes, this telescope is inspired from a smart function of the human eye to improve its availability. This special-made objective lens achieves not only wide-angle surveillance and but also detailed observation on the same time. Moreover, a special-made eyepiece lens part is also designed for users to observe environments with a more immersive feeling. This eyepiece utilizes a broader area on the human retina for the WAF telescope image projection. Indeed, the developed telescope has 22 aspherical surfaces of all 28 surfaces by using economy plastic lenses.
Sota Shimizu, Nobuyuki Hasebe
IECON1
2016 Tessellation for Wide Angle Foveated image with 4 regions based on overlapping circular receptive field mapping
abstract
As well-known, the human single eye has a horizontally approximately 120-degree wide FOV. Its FOV is wide, but only its central FOV has high visual acuity. Visual acuity in its peripheral FOV is very low. By changing its view line into a target, the human can observe it in detail by its central FOV and simultaneously can observe the whole of environment by its wide FOV. This means that the human eye performs a smart data acquisition, i.e., getting more detailed visual information with smaller data amount from environment by combining appropriate eye movements. WAFVS was investigated taking into account the above human eye's functionality. The most characterized point of WAFVS is of the resolution-variant input image data, i.e., this vision sensor gets a wide-angle image where spatial resolution is not uniform like the human visual acuity. Therefore, WAFVS can reduce image data hugely. This is a strong advantage when image data are transmitted and recorded in a storage device when the vision sensor is applied for remote operation. In this paper, a remapping method for WAFVS is proposed and implemented, i.e., tessellation for Wide Angle Foveated image with 4 regions based on overlapping circular RF mapping.
Sota Shimizu, Nobuyuki Hasebe
IECON1
2015 Towards non-mechanical wide angle fovea sensor - Fundamental design by liquid crystal lens cell
abstract
The authors aim at developing a wide angle fovea sensor by which a gaze point and magnification around it are variable in a field of view (FOV) without mechanical parts. Moreover, this sensor can gaze at not only a single point but also multiple points simultaneously. In order to achieve these functions, we apply liquid crystal (LC) for lens materials and control its distribution of optical refraction. As the first step, this paper introduces a prototype of a LC wide-angle fovea lens, in which the magnification around the central FOV of the LC lens cell is controlled electrically as maintaining its image angle of the FOV. In general, as the LC lens cell responds more quickly, it becomes harder to change the distribution of the optical refraction largely. The authors have developed a special-made wide-angle input lens and have combined it with the LC lens cell in order to achieve both the large change and the quick response. In addition, this paper introduces an experimental environment built up for observing projection images and fringe patterns of the LC fovea lens.
Sota Shimizu, Susumu Sato
IECON1
2014 Multi-purpose wide-angle vision system for remote control of planetary exploring rover
abstract
Wide-Angle Fovea Vision Sensor (WAFVS) system was designed and developed being inspired from advantages of the human eye's functions. This system is characterized by its space-variant data acquisition property, i.e., the WAFVS captures a 120-degree wide-angle input image in which its resolution (or magnification) changes like the human visual acuity. As well-known, the human visual acuity is the highest at its central field of view (FOV) and decreases rapidly towards its peripheral FOV. Thus, using the WAFVS, we can observe a target in detail by its central field of view while observing the whole of environment by its wide field of view. In addition, by controlling a view direction of the WAFVS, this WAFVS system gets visual information from the environment more in detail by smaller data amount. Hence, the WAFVS achieves a better performance of data transmission and data storage. One of severe problems in remote control of rovers, UAVs, and satellites is of a pay-load. In this point of view, the authors think that the WAFVS is suitable for the planetary exploring rover because it was originally developed for multi-purpose use of a single vision sensor. This paper describes the multi-purpose use of the WAFVS system, i.e., the following tasks: (1) observing the environment displayed to the operator for the remote navigation of the rover, (2) recording images of important scenes by changing a view direction of the WAFVS, and (3) monitoring if the instruments on the rover work well or not. Moreover, this paper experiments and discusses on how to display images to the operator when an eye-tracking device is applied as a target coordinate input device. Accuracy index, i.e., a measurement error of a target, is defined in order to evaluate performance of a combination among the vision sensor, the coordinate input device and the image display method.
Sota Shimizu, Nobuyuki Hasebe, Kazutaka Nakamura, Hiroki Kusano, Hiroshi Nagaoka, Kyeong J. Kim, Yi Re Choi, Eung Seok Yi
IECON1
2014 A study on color information corrected in human brain - Measurement and evaluation of color propagation
abstract
There exists a well-known visual illusion with respect to color propagation. That is, a human subject confuses a color in the subject's peripheral field of view (FOV) as another color in the central FOV (with a shape of circular disk) spreads into the peripheral FOV or as a color in the peripheral FOV erodes another color in the central FOV (also with a shape of circular disk) when he/she keeps gazing at a specific color visual stimulus in his/her central FOV for more than several seconds. In this paper, the authors experiment this illusion using multiple naive subjects in conditions of the visual stimuli as the central or peripheral colors and the disk size change. This paper has proposed a rate of propagation as a criterion and has analyzed and discussed the illusion using it.
Sota Shimizu, Takumi Kadogawa, Masayuki Naito, Takumi Hashizume, Hiroki Kusano, Hiroshi Nagaoka, Nobuyuki Hasebe, Yoshiaki Tanzawa
IECON1
2012 Development of micro wide angle fovea lens
abstract
This paper describes development of a biologically-inspired wide-angle fovea lens well-known as a significant part of the bio-mimetic vision sensor. The authors develop a micro wide angle fovea (WAF) lens suitable for a board lens camera with a 1/3 inch imaging chip. The prototype of this Micro WAF lens has a smaller diameter and length, i.e., φ8mm × 15mm, than half of the WAF lens the author produced formerly, i.e., φ16mm × 32mm, while its central field of view (FOV), we call fovea, has much higher resolution, i.e., its projected image is more highly-distorted than the former WAF lens. The central FOV is about 10 percent of the entire FOV inside an incident angle of 2.5 degrees in Micro WAF lens, while it is about 10 percent inside an incident angle of 10 degrees in the former WAF lens. We focus mainly on determining its specification as the 1st step of a lens design procedure and discussing results simulated using a lens design software.
Sota Shimizu, Motosuke Kiyohara, Takumi Hashizume
IECON1
2007 Eccentricity Compensator for Log-Polar Sensor
abstract
This paper aims at acquiring robust rotation, scale, and translation-invariant feature from a space-variant image by a fovea sensor. A proposed model of eccentricity compensator corrects deformation that occurs in a log-polar image when the fovea sensor is not centered at a target, that is, when eccentricity exists. An image simulator in discrete space remaps a compensated log-polar image using this model. This paper proposes unreliable feature omission (UFO) that reduces local high frequency noise in the space-variant image using discrete wavelet transform. It discards coefficients when they are regarded as unreliable based on digitized errors of the input image. The first simulation mainly tests geometric performance of the compensator, in case without noise. This result shows the compensator performs well and its root mean square error (RMSE) changes only by up to 2.54 [%] in condition of eccentricity within 34.08[deg]. The second simulation applies UFO to the log-polar image remapped by the compensator, taking its space-variant resolution into account. The result draws a conclusion that UFO performs better in case with more white Gaussian noise (WGN), even if the resolution of the compensated log-polar image is not isotropic.
Sota Shimizu, Joel W. Burdick
ICRA1
2006 Image Extraction by Wide Angle Foveated Lens for Overt-attention
abstract
This paper defines wide angle foveated (WAF) imaging. A proposed model combines Cartesian coordinate system, a log-polar coordinate system, and a unique camera model composed of planar projection and spherical projection for all-purpose use of a single imaging device. The central field-of-view (FOV) and intermediate FOV are given translation-invariance and, rotation and scale-invariance for pattern recognition, respectively. Further, the peripheral FOV is more useful for camera's view direction control, because its image height is linear to an incident angle to the camera model's optical center point. Thus, this imaging model improves its usability especially when a camera is dynamically moved, that is, overt-attention. Moreover, simulation results of image extraction show advantages of the proposed model, in view of its magnification factor of the central FOV, accuracy of scale-invariance and flexibility to describe other WAF vision sensors
Sota Shimizu, Joel W. Burdick
ICRA1
2005 Machine Vision System to Induct Binocular Wide-Angle Foveated Information into Both the Human and Computers - Feature Generation Algorithm based on DFT for Binocular Fixation -
abstract
This paper introduces a machine vision system, which is suitable for cooperative works between the human and computer. This system provides images inputted from a stereo camera head not only to the processor but also to the user’s sight as binocular wide-angle foveated (WAF) information, thus it is applicable for Virtual Reality (VR) systems such as tele-existence or training experts. The stereo camera head plays a role to get required input images foveated by special wide-angle optics under camera view direction control and 3D head mount display (HMD) displays fused 3D images to the user. Moreover, an analog video signal processing device much inspired from a structure of the human visual system realizes a unique way to provide WAF information to plural processors and the user. Therefore, this developed vision system is also much expected to be applicable for the human brain and vision research, because the design concept is to mimic the human visual system. Further, an algorithm to generate features using Discrete Fourier Transform (DFT) for binocular fixation in order to provide well-fused 3D images to 3D HMD is proposed. This paper examines influences of applying this algorithm to space variant images such as WAF images, based on experimental results.
Sota Shimizu, Shinsuke Shimojo, Joel W. Burdick
ICRA1