EDBT 2026 Demo / reviewers in the wild / expert
Jun Ohya
dblp:63/4674
· DBLP profile ↗
77ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-7148-4127ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 38 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DisasterSynth: High-Resolution Disaster Scene Generation and Reliable Pseudo Labeling with Masked Feature Reconstruction
Hiroyuki Ishii, Jun Ohya |
ICPRAM | 3 |
| 2025 | Classification of Oral Cancer and Leukoplakia Using Oral Images and Deep Learning with Multi-Scale Random Crop Self-Training
Itsuki Hamada, Takaaki Ohkawauchi, Chisa Shibayama, Kitaro Yoshimitsu, Nobuyuki Kaibuchi, Katsuhisa Sakaguchi, Toshihiro Okamoto, Jun Ohya |
ICPRAM | 8 |
| 2025 | FFAD: Fixed-Position Few-Shot Anomaly Detection for Wire Harness Utilizing Vision-Language Models
Powei Liao, Pei-Chun Chien, Hiroki Tsukida, Yoichi Kato, Jun Ohya |
ICPRAM | 5 |
| 2025 | A Hierarchical Classification for Automatic Assessment of the Reception Quality Using Videos of Volleyball and Deep Learning
Shota Nako, Hiroyuki Ogata, Taiji Matsui, Itsuki Hamada, Jun Ohya |
ICPRAM | 5 |
| 2025 | A Group Activity Based Method for Early Recognition of Surgical Processes Using the Camera Observing Surgeries in an Operating Room and Spatio-Temporal Graph Based Deep Learning Model
Keishi Nishikawa, Jun Ohya |
ICPRAM | 2 |
| 2025 | A Two-Stage Approach for Wire Harness Cable Description Using 3D Point Clouds for Robotic Manufacturing
Takumi Okuyama, Pei-Chun Chien, Hiroki Tsukida, Yoichi Kato, Jun Ohya |
ICPRAM | 5 |
| 2025 | LAST: Utilizing Synthetic Image Style Transfer to Tackle Domain Shift in Aerial Image Segmentation
Ruijia Wen, Hiroyuki Ishii, Jun Ohya |
ICPRAM | 4 |
| 2024 | Do Text-Free Diffusion Models Learn Discriminative Visual Representations?
Soumik Mukhopadhyay 0001, Matthew Gwilliam, Yosuke Yamaguchi, Vatsal Agarwal, Namitha Padmanabhan, Archana Swaminathan, Tianyi Zhou 0001, Jun Ohya, Abhinav Shrivastava |
ECCV (60) | 8 |
| 2024 | Detecting Overgrown Plant Species Occluding Other Species in Complex Vegetation in Agricultural Fields Based on Temporal Changes in RGB Images and Deep Learning
Haruka Ide, Hiroyuki Ogata, Takuya Otani, Atsuo Takanishi, Jun Ohya |
ICPRAM | 5 |
| 2024 | MAC: Multi-Scales Attention Cascade for Aerial Image Segmentation
Zhao Wang 0009, Yuusuke Nakano, Katsuya Hasegawa, Hiroyuki Ishii, Jun Ohya |
ICPRAM | 6 |
| 2024 | Locating the Fruit to Be Harvested and Estimating Cut Positions from RGBD Images Acquired by a Camera Moved along Fixed Paths Using a Mask-R-CNN Based MethodabstractCompared to traditional agricultural environments, the high density and diversity of vegetation layouts in Synecoculture farms present significant challenges in locating and harvesting occluded fruits and pedicels (cutting points). To address this challenge, this study proposes a Mask R-CNN-based method for locating fruits (tomatoes, yellow bell peppers, etc.) and estimating the pedicels from RGBD images acquired by a camera moved along fixed paths. After obtaining masks of all fruits and pedicels, this method judges the matching relationship between the located fruit and pedicel according to the 3D distance between the fruit and pedicel. Subsequently, this research determines the least occluded best viewpoint for harvesting based on the visible real areas of located fruits in images acquired under the fixed paths, and harvesting is then completed from this best viewpoint following a straight path. Experimental results show this method effectively identifies occluded targets and their cutting positions in both Gazebo simulation environments and real-world farms. This method can select the least occluded viewpoint for a high harvesting success rate. Takuya Otani, Sugiyama Soma, Mitani Kento, Koki Masaya, Atsuo Takanishi, Shuntaro Aotake, Masatoshi Funabashi, Jun Ohya |
RO-MAN | 9 |
| 2024 | Classifying Cable Tendency with Semantic Segmentation by Utilizing Real and Simulated RGB DataabstractCable tendency is the potential shape or characteristic that a cable may possess while being manipulated, of which some are considered erroneous and should be identified as a part of anomaly detection during an automatic manipulation. This research explores the ability of deep-learning models in learning the cable tendencies that, contrary to typical classification tasks of multi-object scenarios, is to differentiate the multiple states displayable by the same object – in this case, cables. By training multiple models with different combinations of self-collected real-world data and self-generated simulation data, a comparative study is carried out to compare the performance of each approach. In conclusion, the effectiveness of detecting three abnormal states and shapes of cables, and using simulation data is certificated in experiments. Pei-Chun Chien, Powei Liao, Eiji Fukuzawa, Jun Ohya |
WACV | 4 |
| 2023 | Virtual Ski Training System that Allows Beginners to Acquire Ski Skills Based on Physical and Visual FeedbacksabstractThis paper proposes a ski training system using VR (Virtual Reality) that enables beginners to acquire skiing skills without going to an actual ski ground. The proposed system obtains the speed of skiing based on the center of pressure (COP) of each player's foot. The first-person perspective of skiing at the obtained speed down a ski slope is fed back to the player as a VR image. Experiments were conducted to evaluate the effectiveness of the proposed system and the VR interface. Specifically, beginner skiers were categorized into three groups: “a group trained with the proposed VR system”, “a group trained with a system that provides feedback of the skiing speed calculated from the COP by increasing or decreasing the gauge (a bar-shaped graph representing changes in numerical values), instead of VR”, and “a group that does not train with the system”. After training under each of these conditions, a sliding test was conducted on an actual ski slope to check the degree of skill acquisition. The results show that subjects trained with the proposed system acquired more skiing skills than subjects who did not use the system on actual ski slopes. Furthermore, there was no clear difference in the result of the sliding test between subjects trained by the VR interface and those trained by the gauge interface, but the VR interface yields better deceleration postures. Yushi Okada, Chanjin Seo, Shunichi Miyakawa, Motofumi Taniguchi, Kazuyuki Kanosue, Hiroyuki Ogata, Jun Ohya |
IROS | 7 |
| 2022 | Preliminary Investigation of Collision Risk Assessment with Vision for Selecting Targets Paid Attention to by Mobile RobotabstractVision plays an important role in motion planning for mobile robots which coexist with humans. Because a method predicting a pedestrian path with a camera has a trade-off relationship between the calculation speed and accuracy, such a path prediction method is not good at instantaneously detecting multiple people at a distance. In this study, we thus present a method with visual recognition and prediction of transition of human action states to assess the risk of collision for selecting the avoidance target. The proposed system calculates the risk assessment score based on recognition of human body direction, human walking patterns with an object, and face orientation as well as prediction of transition of human action states. First, we investigated the validation of each recognition model, and we confirmed that the proposed system can recognize and predict human actions with high accuracy ahead of 3 m. Then, we compared the risk assessment score with video interviews to ask a human whom a mobile robot should pay attention to, and we found that the proposed system could capture the features of human states that people pay attention to when avoiding collision with other people from vision. Masaaki Hayashi, Tamon Miyake, Mitsuhiro Kamezaki, Junji Yamato, Kyosuke Saito, Taro Hamada, Eriko Sakurai, Shigeki Sugano, Jun Ohya |
RO-MAN | 9 |
| 2021 | Quantitative Method for Evaluating the Coordination between Sprinting Motions using Joint Coordinates Obtained from the Videos and Cross-correlations
Masato Sabanai, Chanjin Seo, Hiroyuki Ogata, Jun Ohya |
ICPRAM | 4 |
| 2021 | Movement Control with Vehicle-to-Vehicle Communication by using End-to-End Deep Learning for Autonomous Driving
Jun Ohya |
ICPRAM | 2 |
| 2020 | Extracting and Interpreting Unknown Factors with Classifier for Foot Strike Types in RunningabstractThis paper proposes a method that can classify foot strike types using a deep learning model and can extract unknown factors, which enables to evaluate running motions without being influenced by biases of sports experts, using the contribution degree of input values (CDIV). Accelerometers are attached to the runner's body, and when the runner runs, a fixed camera observes the runner and acquires a video sequence synchronously with the accelerometers. To train a deep learning model for classifying foot strikes, we annotate foot strike acceleration data for RFS (Rearfoot strike) or non-RFS objectively by watching the video. To interpret the unknown factors extracted from the learned model, we calculate two CDIVs: the contributions of the resampling time and the accelerometer value to the output (foot strike type). Experiments on classifying unknown runners' foot strikes were conducted. As a common result to sport science, it is confirmed that the CDIVs contribute highly at the time of the right foot strike, and the sensor values corresponding to the right and left tibias contribute highly to classifying the foot strikes. Experimental results show the right tibia is important for classifying foot strikes. This is because many of the training data represent difference between the two foot strikes in the right tibia. As a conclusion, our proposed method could extract unknown factors from the classifier and could interpret the factors that contain similar knowledge to the prior knowledge of experts, as well as new findings that are not included in conventional knowledge. Chanjin Seo, Masato Sabanai, Yuta Goto, Koji Tagami, Hiroyuki Ogata, Kazuyuki Kanosue, Jun Ohya |
ICPR | 7 |
| 2020 | Developing Thermal Endoscope for Endoscopic Photothermal Therapy for Peritoneal DisseminationabstractAs a novel therapy for peritoneal dissemination, it is desired to actualize an endoscopic photothermal therapy, which is minimally invasive and is highly therapeutically effective. However, since the endoscopic tumor temperature control has not been actualized, conventional therapies could damage healthy tissues by overhearing. In this paper, we develop a thermal endoscope system that controls the tumor temperature so that the heated tumor gets necrotic. In fact, our thermal endoscope contains a thermal image sensor, a visible light endoscope and a laser fiber. Concerning the thermal image sensor, the conventional thermal endoscope has the problem that the diameter is too large, because the conventional endoscope loads a large thermal image sensor with high-resolution. Therefore, this paper uses a small thermal image sensor with low resolution, because the diameter of the thermal endoscope needs to be smaller than 15mm in order to be inserted into the trocar. However, this thermal image sensor is contaminated by much noise. Thus, we develop a tumor temperature control system using a feedback control and tumor temperature estimation based on Gaussian function, so that the noisy, small thermal image sensor can be used. As experimental results of the proposed endoscopic photothermal therapy for the hepatophyma carcinoma model of rats, it turns out that the tumor temperature by which the heated tumor gets necrotic can be kept stable. It can be said that our endoscopic photothermal therapy achieves a certain degree of therapy effect. Mutsuki Ohara, Sohta Sanpei, Chanjin Seo, Jun Ohya, Ken Masamune, Hiroshi Nagahashi, Yuji Morimoto, Manabu Harada |
IROS | 4 |
| 2019 | Detecting and Tracking Surgical Tools for Recognizing Phases of the Awake Brain Tumor Removal SurgeryabstractIn order to realize automatic recognition of surgical processes in surgical brain tumor removal using microscopic camera, we propose a method of detecting and tracking surgical tools by video analysis. The proposed method consists of a detection part and tracking part. In the detection part, object detection is performed for each frame of surgery video, and the category and bounding box are acquired frame by frame. The convolution layer strengthens the robustness using data augmentation (central cropping and random erasing). The tracking part uses SORT, which predicts and updates the acquired bounding box corrected by using Kalman Filter; next, the object ID is assigned to each corrected bounding box using the Hungarian algorithm. The accuracy of our proposed method is very high as follows. As a result of experiments on spatial detection. the mean average precision is 90.58%. the mean accuracy of frame label detection is 96.58%. These results are very promising for surgical phase recognition. Hiroki Fujie, Keiju Hirata, Takahiro Horigome, Hiroshi Nagahashi, Jun Ohya, Manabu Tamura, Ken Masamune, Yoshihiro Muragaki |
ICPRAM | 5 |
| 2019 | Detecting a Fetus in Ultrasound Images using Grad CAM and Locating the Fetus in the UterusabstractIn this paper, we propose an automatic method for estimating fetal position based on classification and detection of different fetal parts in ultrasound images. Fine tuning is performed in the ultrasound images to be used for fetal examination using CNN, and classification of four classes leg and other is realized. Based on the obtained learning result, binarization that thresholds the gradient of the feature obtained by Grad Cam is performed in the image so that a bounding box of the region of interest with large gradient is extracted. The center of the bounding box is obtained from each frame so that the trajectory of the centroids is obtained; the position of the fetus is obtained as the trajectory. Experiments using 2000 images were conducted using a fetal phantom. Each recall ratiso of the four class is 99.6% for head, 99.4% for body, 99.8% for legs, 72.6% for others, respectively. The trajectories obtained from the fetus present in “left”, “center”, “right” in the images show the above-mentioned geometrical relationship. These results indicate that the estimated fetal position coincides with the actual position very well, which can be used as the first step for automatic fetal examination by robotic systems. Genta Ishikawa, Jun Ohya, Hiroyasu Iwata |
ICPRAM | 3 |
| 2019 | Understanding Sprinting Motion Skills using Unsupervised Learning for Stepwise Skill Improvements of Running MotionabstractTo improve running performances, each runner’s skill, such as characteristics and habits, needs to be known, and feedback on the performance should be outputted according to the runner's skill level. In this paper, we propose a new coaching system for detecting the skill of a runner and a method of giving feedback using a sprint motion dataset. Our proposed method calculates an extracted feature to detect the skill using an autoencoder whose middle layer is an LSTM layer; we analyse the feature using hierarchical clustering, and we analyse the human joints that affect the skill. As a result of experiments, five clusters are obtained using hierarchical clustering. This paper clarifies how to detect the skill and to output feedback to achieve a level of performance one step higher than the current level. Chanjin Seo, Masato Sabanai, Hiroyuki Ogata, Jun Ohya |
ICPRAM | 4 |
| 2019 | Disaster Response Robot's Autonomous Manipulation of Valves in Disaster Sites Based on Visual Analyses of RGBD ImagesabstractFor building a disaster response robot, WAREC-l's fully-automated system for manipulating a valve, this paper proposes a method for (1) detecting a valve which is far away from the robot, (2) estimating the position and orientation for grasping the valve by the robot at a closer position. Our methods do not need any prior information about a valve for the above-mentioned detection and estimation for grasping. In addition, our estimation for grasping provides useful information, by which WAREC-I can rotate a valve autonomously. The method (1) uses the RGB image and the point cloud data captured by Multisense SL as the input, and estimate the position and orientation of a valve far away from the robot. The method (2) uses both the RGB and depth images captured by KinectV2 as input and estimate information for grasping the valve. Our experiments are conducted using a real disaster response robot. Our experimental results show the error of the estimation by the (a) two methods are small enough to achieve a fully-automated system for detecting and rotating the valve by WAREC-I. Keishi Nishikawa, Asaki Imai, Kazuya Miyakawa, Takuya Kanda, Takashi Matsuzawa, Kenji Hashimoto, Atsuo Takanishi, Hiroyuki Ogata, Jun Ohya |
IROS | 9 |
| 2017 | Exploring the effectiveness of using temporal order information for the early-recognition of suture surgery's six steps based on video image analyses of surgeons' hand actionsabstractTo alleviate the recent shortage problem of nurses, the actualization of RSN (Robotic Scrub Nurse) that can autonomously judge the current step of the surgery and pass the surgical instruments needed for the next step to surgeons is desired. The authors developed a computer vision based algorithm that can early-recognize only two steps of suture surgery. Based on the past work, this paper explores the effectiveness of utilizing temporal order of the six steps in suture surgery for the early-recognition. Our early-recognition algorithm consists of two modules: start point detection and hand action early-recognition. Segments of the test video that start from each quasi-start point are compared with the training data, and their probabilities are calculated. According to the calculated probabilities, hand actions could be early-recognized. To improve the early-recognition accuracy, temporal order information could be useful. This paper checks confusions of three steps' early recognition results, and if necessary, early-recognizes again after eliminating the wrong result, while for the other three steps, temporal order information is not utilized. Experimental results show our early-recognition method that utilizes the temporal order information achieves better performances. Miwa Tsubota, Jun Ohya |
RO-MAN | 3 |
| 2010 | Study on human gesture recognition from moving camera imagesabstractWe develop a framework based approach to extract and recognize hand gestures from the video sequence acquired by a dynamic camera, which could be a useful interface between humans and mobile robots. We use Human-Following Local Coordinate (HFLC) System, a very simple and stable method for extracting hand motion trajectories, which is obtained from the located human face and body part. Hand trajectory motion models (HTMM) are constructed by HFLC and hand blob changing factor. In this paper, we apply a principal component analysis (PCA) based approach to improve the recognition accuracy. For further improvement, temporal changes in the observed hand area changing factor are utilized as new image features to be stored in the database after being analyzed by PCA. Each HTMM in the database is classified into gesture categories, or temporal changes in hand blob changes. We demonstrate the effectiveness of the proposed method by conducting experiments on 51 kinds of sign language based Japanese and American Sign Language gestures obtained from 7 people. Our experimental recognition results show better performance is obtained by PCA based approach than the Condensation algorithm based method. Dan Luo 0002, Jun Ohya |
ICME | 2 |
| 2007 | A study of a computer mediated communication via the "." prompt system - Analysis of the affects on the stimulation of thought processes and the inspiration of creative ideasabstractResearch into thinking-support tools is commonly focused on how to develop and share ideas between participants or with others. In this paper, we propose and develop a communication system that stimulates the thought processes and inspires the creative ideas of participants by using a visual "" prompt within the framework of a communication pallet. Experiments have been conducted into methods of stimulating the thought process and inspiring ideas during conversation and the results have been analyzed. From the results, a tendency towards inspiring creative ideas by participants has been observed. Li Jen Chen, Nobuyuki Harada, Shunichi Yonemura, Jun Ohya, Yukio Tokunaga |
ICME | 4 |
| 2005 | A Study of Synthesizing New Human Motions from Sampled Motions Using Tensor DecompositionabstractThis paper applies an algorithm, based on Tensor Decom position, to a new synthesis application: by using sampled motions of people of different ages under different emotional states, new motions for other people are synthesized. Human motion is the composite consequence of multiple elements, including the action performed and a motion signature that captures the distinctive pattern of movement of a particular individual. By performing decomposition, based on N-mode SVD (singular value decomposition), the algorithm analyzes motion data spanning multiple subjects performing different actions to extract these motion elements. The analysis yields a generative motion model that can synthesize new motions in the distinctive styles of these individuals. The effectiveness of applying the tensor decomposition approach to our purpose was confirmed by synthesizing novel walking motions for a person by using the extracted signature. Rovshan Kalanov, Jieun Cho, Jun Ohya |
ICME | 3 |
| 2005 | Analysis of expressing audiences in a cyber-theaterabstractThis paper studies how audiences should be expressed in a Cyber-theater, in which remotely located persons can direct plays as directors, perform as performers and/or see the performances as audiences through a networked virtual environment. It is noted that the audience effect has been widely acknowledged in the real-world theater: that is, the audience reaction has a significant effect on the acting of player and performance of the play itself. However, only a few works relevant to audiences in the cyber theater can be seen. This paper studies whether the audience effect exists also in the cyber-theater. By constructing, a system in which two actors are displayed a remotely located audience's avatar in which the audience can display his/her emotional actions, we clarified that interaction between the actors and audiences are effective. Don Wan Kang, Kay Huang, Jun Ohya |
ICME | 3 |
| 2004 | Visual-dimension Interact System (VIS) - Exhibiting Creative Process for Museum Visitor Experience -abstractWe describe a mixed reality-supported interactive viewing enhancement museum display system. With a transparent interactive interface, the museum visitor is able to see, manipulate, and interact with the physical exhibit and its virtual information, which are overlapped on one other. Furthermore, this system provides the possibility for the visitor to experience the creation process in an environment as close as possible to the real process. This has the function of assisting the viewer in understanding the exhibit and most importantly, gaining a so-to-speak hands-on experience of the creation process itself leading to a deeper understanding of it. Atsushi Onda, Tomoyuki Oku, Pei-Yi Chiu, Eddie Yu, Maki Yokoi, Ikuro Choh, Jun Ohya |
CW | 7 |
| 2004 | Cognitive bridge between haptic impressions and texture images for subjective image retrievalabstractAs a step towards subjective image retrieval, This work reports an on-going collaboration project between Waseda University and SUNY, Binghamton, on relating texture images to haptic impressions. To capture the surface height variations, texture images are taken under different illuminations and viewing conditions. Our method applies a new frequency analysis method to the texture images. We evaluate the performances of our feature and other typical conventional features by checking whether texture images are correctly classified into "soft" or "hard" by the SVM (support vector machine) method, where the training data for the SVM are collected by subjective tests. Experimental results show that our texture feature can classify "soft" or "hard" better than the other features. Yuichi Kobayashi, Jun Ohya, Zhongfei Zhang |
ICME | 2 |
| 2004 | Computer vision based analysis of the botanical tree's dynamical behaviors for the reproduction in virtual spaceabstractThe paper deals with a method that analyzes a botanical tree's behaviors in real space by a computer vision approach so as to reproduce the analyzed behavior in virtual space. Instead of applying unstable local tracking to the tree in a video sequence, we estimate the direction and strength of the wind that shakes the tree by a learning based method that classifies the input video sequence into one of the stored winds with different directions and strengths. In the learning phase, sample video sequences are used for constructing the eigenspace and Fisherspace, which is obtained from Fisher discriminant analysis. In the classification phase, the input video sequence is compared with each of the stored sample sequences so that the direction and strength of the wind are estimated. An interpolation method improves the estimation accuracy. Experimental results demonstrate the effectiveness of the proposed method. Liang-Chen Lu, Jun Ohya |
ICME | 2 |
| 2004 | Exploiting the cognitive synergy between different media modalities in multimodal information retrievalabstractThis is a position paper reporting an on-going collaboration project between SUNY Binghamton, USA, and Waseda University, Japan, on multimodal information retrieval through exploiting the cognitive synergy across the different modalities of the information, to facilitate an effective retrieval. Specifically, we focus on image retrieval in the applications where imagery data appear along with collateral text. It is noted that these applications are ubiquitous. We have proposed the synergistic indexing scheme (SIS) to explicitly exploit the synergy between the information of imagery and text modalities. Since the synergy we have exploited between the information of imagery and text modalities is subjective and depends on specific cognitive context, we call this type of synergy as cognitive synergy. We have reported part of the empirical evaluation and are in the process of fully implementing the SIS prototype for an extensive evaluation. Zhongfei Zhang, Ruofei Zhang, Jun Ohya |
ICME | 3 |
| 2003 | Efficient, realistic method for animating dynamic behaviors of 3D botanical treesabstractThis paper proposes a new efficient method that can animate botanical trees in 3D realistically. In this paper, a 3D botanical tree model consists of a set of branch segments, to which leaf models are attached. To reduce the amount of computation, instead of calculating the motions of all the branch segments, only the representative segment in each branch is numerically analyzed. The numerical analysis is constrained to a 2D plane so that 3D numerical analysis need not be performed. Concerning the leaf model, a set of four leaves is systematically attached to each branch segment. Experimental results clarify the conditions for real-time, realistic animations of dynamic behaviors of trees. Hitoshi Kanda, Jun Ohya |
ICME | 2 |
| 2003 | Postures of a human wearing a multiple-colored suit based on color information processingabstractThis paper proposes a non-contact type method for estimating human body postures. One of the major problems on the posture estimation using the silhouette image analysis is the overlapping of the body parts' silhouettes. In order to solve this problem, this paper proposes a method for estimating the posture of a human wearing a multiple-colored suit based on color information processing. By analyzing the contour of the human's silhouette, the method judges whether feature points are occluded by another body parts. If the occlusions occurs, color region segmentation is performed in order to know which region is frontal. The feature point in the frontal region is located in the skeleton of the region. Experimental results show the effectiveness of the proposed method. Dong-Wan Kang, Jun Ohya |
ICME | 2 |
| 2002 | Construction of facial expressions using a muscle-based feature modelabstractAn efficient method for constructing facial images for use in telecommunication applications is proposed. This method uses a simple 3D feature model, which consists of polygons, which describe the shape of the face, and elastic linear springs, which simulate the natural movements of facial muscles. This method requires only two orthogonal facial images, and could easily be implemented on a relatively low-spec PC. Experimental results showed good results that various facial expressions could be synthesized and displayed from arbitrary directions. Yi-Chih Liu, Hajime Sato, Nobuyoshi Terashima, Jun Ohya |
ICME (1) | 4 |
| 2002 | Analysis of human behaviors by computer vision based approachesabstractThis paper describes the author's activities related to computer vision based methods for analyzing human behaviors: more specifically, posture estimation and recognizing interactions between a human body and object. For estimating postures in 3D from multiple camera images, the authors developed a heuristic based method and non-heuristic method. The heuristic based method heuristically analyzes the contour of a human silhouette so that significant points of a human body can be located in each image. The non-heuristic method utilizes a function for analyzing contours without using heuristic rules. Recognizing the interactions exploits the function based contour analysis and motion vector based analysis so that the system can judge whether the human body interacts with the object. Jun Ohya |
ICME (1) | 1 |
| 2002 | Face posture estimation using eigen analysis on an IBR (image based rendered) database
Kuntal Sengupta, Philip Lee, Jun Ohya |
Pattern Recognit. | 3 |
| 2001 | Computer Vision Based Analysis of Non-verbal Information in HCIabstractThis paper overviews our research activities on computer vision based non-verbal information analysis that can be applied to virtual communication environments and human computer interactions. In virtual communication environments, a user’s facial expressions and body motions are estimated by computer vision approaches, and the estimated non-verbal information is reproduced in the user’s avatar. For human computer interfaces, hand gestures are recognized as pre-defined commands by analyzing multiple camera images that observe the hand. In addition, facial expressions and body gestures are recognized from a time-sequential images by HMM (Hidden Markov Models). Jun Ohya |
ICME | 1 |
| 2001 | User-guided composition effects for art-based renderingabstractWe apply simple techniques from traditional artistic composition to the art-based rendering of interactive 3D scenes. A human scene-modeler makes choices about composition in a scene and our sys-tem dynamically adjusts the rendering attributes of objects in the scene to achieve the desired effects for a given view. We can selec-tively group scene elements through shared tone, color, and outline, so as to simplify and structure an image. This can be used, together with controlled level of detail, to emphasize important objects. Fi-nally, we show a technique for adaptively changing color or other attributes to control the contrast of adjacent elements in the pic-ture. We also briefly discuss ideas about larger-scale compositional issues. Michael A. Kowalski, John F. Hughes, Cynthia Beth Rubin, Jun Ohya |
SI3D | 4 |
| 2001 | Spatial Filtering Using the Active-Space Indexing Method
Sudhanshu Kumar Semwal, Jun Ohya |
Graph. Model. | 2 |
| 2000 | Human Body Postures from Trinocular Camera ImagesabstractThis paper proposes a new real-time method for estimating human postures in 3D from trinocular images. In this method, an upper body orientation detection and a heuristic contour analysis are performed on the human silhouettes extracted from the trinocular images so that representative points such as the top of the head can be located. The major joint positions are estimated based on a genetic algorithm-based learning procedure. 3D coordinates of the representative points and joints are then obtained from the two views by evaluating the appropriateness of the three views. The proposed method implemented on a personal computer runs in real-time. Experimental results show high estimation accuracies and the effectiveness of the view selection process. Shoichiro Iwasawa, Jun Ohya, Kazuhiko Takahashi, Tatsumi Sakaguchi, Shigeo Morishima, Kazuyuki Ebihara |
FG | 2 |
| 2000 | Real-Time Detection of Nodding and Head-Shaking by Directly Detecting and Tracking the "Between-Eyes"abstractAmong head gestures, nodding and head-shaking are very common and used often. Thus the detection of such gestures is basic to a visual understanding of human responses. However it is difficult to detect them in real-time, because nodding and head-shaking are fairly small and fast head movements. We propose an approach for detecting nodding and head-shaking in real time from a single color video stream by directly detecting and tracking a point between the eyes, or what we call the "between-eyes". Along a circle of a certain radius centered at the "between-eyes", the pixel value has two cycles of bright parts (forehead and nose bridge) and dark parts (eyes and brows). The output of the proposed circle-frequency filter has a local maximum at these characteristic points. To distinguish the true "between-eyes" from similar characteristic points in other face parts, we do a confirmation with eye detection. Once the "between-eyes" is detected, a small area around it is copied as a template and the system enters the tracking mode. Combining with the circle-frequency filtering and the template, the tracking is done not by searching around but by selecting candidates using the template; the template is then updated. Due to this special tracking algorithm, the system can track the "between-eyes" stably and accurately. It runs at 13 frames/s rate without special hardware. By analyzing the movement of the point, we can detect nodding and head-shaking. Some experimental results are shown. Shinjiro Kawato, Jun Ohya |
FG | 2 |
| 2000 | Remarks on a Real-Time 3D Human Body Posture Estimation Method Using Trinocular ImagesabstractThis paper proposes a new real-time method of estimating human postures in 3D form trinocular images. The proposed method extracts feature points of the human body by applying a type of function analysis to contours of human silhouettes. To overcome self-occlusion problems, dynamic compensation is carried out using the Kalman filter and all feature points are tracked. The 3D coordinates of the feature points are reconstructed by considering the geometrical relationship between the three cameras. Experimental results confirm both the feasibility and the effectiveness of the proposed method, and an application example of the 3D human body posture estimation to a motion recognition system is presented. Kazuhiko Takahashi, Tatsumi Sakaguchi, Jun Ohya |
ICPR | 3 |
| 2000 | Adaptive Human Motion Tracking Using Non-Synchronous Multiple Viewpoint ObservationsabstractWe propose an adaptive human tracking system with non-synchronous multiple observations. Our system consists of three types of processes: discovering node for detecting newly appeared person; tracking node for tracking each target person; and observation node for processing one viewpoint (camera) images. We have multiple observation nodes and each node works independently. The tracking node integrates the observed information based on reliability evaluation. Both the observation conditions and human motion states are considered in the evaluation. Matching between tracking models and observed image features are performed in each of the observation node based on the position, size and color similarities of each 2D image. Due to the non-synchronous property, this system is highly scalable for increasing the detection area and number of observing nodes. Experimental results for some indoor scenes are also described. Akira Utsumi, Howard Yang, Jun Ohya |
ICPR | 3 |
| 2000 | Two-step approach for real-time eye tracking with a new filtering techniqueabstractHead and face detection and eye tracking in real time are the first steps for head gesture recognition and/or face expression recognition for a human-computer interaction interface. We propose a two-step approach for eye tracking in video streams. First, we detect or track a point between the eyes. For this task, we apply a special filter that we have proposed previously (Proc. IEEE 4th Int. Conf. on Automatic Face and Gesture Recognition, pp. 40-45, 2000). Once we have detected the point between the eyes, it is fairly easy to locate the eyes, which are the two small darkest parts on each side of this point. Because detecting a point between the eyes is easier and more stable than detecting the eyes directly, the system can track the eyes robustly. We implemented the system on an SGI O2 workstation. The video image size is 320/spl times/240 pixels. The system processes images at seven frames per second in the detection mode and at 13 frames per second in the tracking mode, without any special hardware. Shinjiro Kawato, Jun Ohya |
SMC | 2 |
| 1999 | Multiple-Hand-Gesture Tracking using Multiple CamerasabstractWe propose a method of tracking 3D position, posture, and shapes of human hands from multiple-viewpoint images. Self-occlusion and hand-hand occlusion are serious problems in the vision-based hand tracking. Our system employs multiple-viewpoint and viewpoint selection mechanism to reduce these problems. Each hand position is tracked with a Kalman filler and the motion vectors are updated with image features in selected images that do not include hand-hand occlusion. 3D hand postures are estimated with a small number of reliable image features. These features are extracted based on distance transformation, and they are robust against changes in hand shape and self-occlusion. Finally, a "best view" image is selected for each hand for shape recognition. The shape recognition process is based on a Fourier descriptor. Our system can be used as a user interface device an a virtual environment, replacing glove-type devices and overcoming most of the disadvantages of contact-type devices. Akira Utsumi, Jun Ohya |
CVPR | 2 |
| 1999 | Modeling and animation of botanical trees for interactive virtual environmentsabstractThis paper proposes a new modeling and animation method of botanical tree for interactive virtual environment. Some studies of botanical tree modeling have been based on the Growth Model, which can construct a very natural tree structure. However, this model makes it difficult to predict the final form of tree from given parameters; that is, if an objective form of a tree is given and it is to be reconstructed into a three-dimensional model, we have to change the parameters to reflect the structure by a trial-and-error technique. Thus, we propose a new top-down approach in which a tree's form is defined by volume data that is made from a captured real image set, and the branch structure is realized by simple branching rules. The tree model is described as a set of connected branch segments, and leaf models that consist of leaves and twigs that are attached to the branch segments. To animate the botanical trees, dynamics simulation is performed on the branch segments in two phases. In the first phase, each segment is assumed to be a rigid stick with a fixed end on one side, and rotational movements from influence of external forces are calculated in each segment independently. And the forces propagated from the tip of a branch to the root are calculated from the restoration force and thickness of the branch. Finally, the rotational movements of segments are executed in order from the base segment, and the fixed end of each segment is moved to the free end of the segment to be connected so as to maintain the relative angles between the segments. The proposed model is applied to many kind of botanical trees, and the model can successfully animate tree movements caused by external forces such as winds and human interaction to the branches. Tatsumi Sakaguchi, Jun Ohya |
VRST | 2 |
| 1998 | Converting Facial Expressions Using Recognition-Based Analysis of Image Sequences
Takahiro Otsuka, Jun Ohya |
ACCV (2) | 2 |
| 1998 | Multiple Camera Based Human Motion Estimation
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
ACCV (2) | 3 |
| 1998 | Recognizing Abruptly Changing Facial Expressions from Time-Sequential Face ImagesabstractThis paper proposes a method that can spot and recognize each facial expression. From time-sequential images that contain multiple facial expressions that could abruptly change from one expression to another expression. Previously, the authors have proposed an HMM (Hidden Markov Models) based method for recognizing a spotted facial expression. In this paper to HMM, we add states corresponding to the simultaneous motion of two different facial expressions: i.e. a muscle relaxation for one expression and a muscle contraction for another expression. Then, the added states are each lined from the HMM apex state of one expression and are linked to that of another expression. Experimental results showed that for most pairs of expressions the change in expression can be recognized accurately. In addition, recognition rate for very fast change of expressions improved significantly. The proposed method was applied to regenerate facial expressions on a synthesized character to show the method's effectiveness in obtaining facial motion information. Takahiro Otsuka, Jun Ohya |
CVPR | 2 |
| 1998 | Image Segmentation for Human Tracking Using Sequential-Image-Based Hierarchical AdaptationabstractWe propose a novel method of extracting a moving object region from each frame in a series of images regardless of complex, changing background using statistical knowledge about the target. In vision systems for 'real worlds' like a human motion tracer, a priori knowledge about the target and environment is often limited (e.g., only the approximate size of the target is known) and is insufficient for extracting the target motion directly. In our approach, information about both target object and environment is extracted with a small amount of given knowledge about the target object. Pixel value (color, intensity, etc.) distributions for both the target object and background region are adaptively estimated from the input image sequence based on the knowledge. Then, the probability of each pixel being associated with the target object is calculated. The target motion can be extracted from the calculated stochastic image. We confirmed the stability of this approach through experiments. Akira Utsumi, Jun Ohya |
CVPR | 2 |
| 1998 | Real-Time Human Posture Estimation Using Monocular Thermal Images
Shoichiro Iwasawa, Kazuyuki Ebihara, Jun Ohya, Shigeo Morishima |
FG | 3 |
| 1998 | Spotting Segments Displaying Facial Expression from Image Sequences Using HMM
Takahiro Otsuka, Jun Ohya |
FG | 2 |
| 1998 | Geometric-Imprints: A Significant Points Extraction Method for the Scan&Track Virtual Environment
Sudhanshu Kumar Semwal, Jun Ohya |
FG | 2 |
| 1998 | Human Face Structure Estimation from Multiple Images Using the 2D Affine Space
Kuntal Sengupta, Jun Ohya |
FG | 2 |
| 1998 | Multiple-Human Tracking Using Multiple Cameras
Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
FG | 3 |
| 1998 | A New Robust Real-Time Method for Extracting Human Silhouettes from Color Images
Masanori Yamada, Kazuyuki Ebihara, Jun Ohya |
FG | 3 |
| 1998 | A new camera projection model and its application in reprojectionabstractWe present a new camera projection model, which is intermediate between the affine camera model and the pin hole projection model. It is modeled as a perspective projection of 3D points into an arbitrary plane, followed by an affine transform of these projected points. We observe that the reprojection of a point into a novel image can be achieved uniquely provided that we have located a set of five reference points over four images (of which three are input images, and the fourth is the novel image). Also, the reprojection theory does not assume that the input images are captured from cameras with identical internal calibration parameters. Thus, we apply our technique to two different domains: 1) generation of novel images from a stereo pair; and 2) generation of virtual walk-through sequence with a monocular image sequence as input. Kuntal Sengupta, Jun Ohya |
ICPR | 2 |
| 1998 | Multiple-view-based tracking of multiple humansabstractWe propose a multiple-view-based tracking algorithm for multiple-human motions. In vision-based human tracking, self-occlusions and human-human occlusions are a part of the more significant problems. Employing multiple viewpoints and a viewpoint selection mechanism, however can reduce these problems. In our system, human positions are tracked with a sequence of multiple-viewpoint images. This tracking is based on the Kalman filtering approach. The estimation results are utilized to select proper viewpoints in other sub-tasks (rotation angle detection and body-side detection). Each sub-task has a different criterion for selecting viewpoints. We also describe the criterions for accomplishing individual sub-tasks and relationships between sub-tasks. We have already built an experimental system based on a small number of reliable image features. We confirm the stability of our algorithm through simulations. We also performed fundamental examinations on the experimental system. Akira Utsumi, Hiroki Mori, Jun Ohya, Masahiko Yachida |
ICPR | 3 |
| 1998 | Real-time estimation of head motion using weak perspective epipolar geometryabstractFor face and facial expression recognition, it is necessary to estimate head motion in order to track a head continuously. This paper proposes a new method for estimating head motion using the epipolar geometry of a weak perspective projection model. In this method, first, the head region is segmented from the gradient of a luminance, by approximating the contour of the head as a circle. Then, feature points such as local extremum or saddle points of a luminance distribution are traded over successive frames. Finally, angles of rotation of the head during two successive frames are estimated from the coordinates of those feature points successfully tracked in these frames. Experiments were performed on a workstation in real time and the results showed that the method performs well in estimating head motion. Takahiro Otsuka, Jun Ohya |
WACV | 2 |
| 1998 | Novel scene generation, merging and stitching views using the 2D affine space
Kuntal Sengupta, Jun Ohya |
Signal Process. Image Commun. | 2 |
| 1997 | Real-Time Estimation of Human Body Posture from Monocular Thermal ImagesabstractThis paper introduces a new real-time method to estimate the posture of a human from thermal images acquired by an infrared camera regardless of the back-ground and lighting conditions. Distance transformation is performed for the human body area extracted from the thresholded thermal image for the. Calculation of the center of gravity. After the orientation of the upper half of the body is obtained by calculating the moment of inertia, significant points such as the top of the head, the tips of the hands and foot are heuristically located. In addition, the elbow and foot positions are estimated from the detected (significant) points using a genetic algorithm based learning procedure. The experimental results demonstrate the robustness of the proposed algorithm and real-time (faster than 20 frames per second) performance. Shoichiro Iwasawa, Kazuyuki Ebihara, Jun Ohya, Shigeo Morishima |
CVPR | 3 |
| 1997 | Recognizing Multiple Persons' Facial Expressions Using HMM Based on Automatic Extraction of Significant Frames from Image SequencesabstractA method that can be used for recognizing facial expressions of multiple persons is proposed. In this method, the condition of facial muscles is assigned to a hidden state of a HMM for each expression. Then, the probability of the state is updated according to a feature vector obtained from image processing. Image processing is performed in two steps. First, a velocity vector is estimated from every two successive frames by using an optical flow algorithm. Then, a two-dimensional Fourier transform is applied to a velocity vector field at the regions around an eye and the mouth. The coefficients for lower frequencies are selected to form a feature vector. A mixture density is used for approximating the output probability of the HMM so as to represent a variation in facial expressions among persons. To cope with the case when two expressions are displayed continuously, the HMM computation is modified such that when the peak of a facial motion is detected, a new sequence of facial expressions is assumed to start from the previous frame with minimal facial motion. Experiments show that a mixture density is effective because the recognition accuracy improves as the number of mixtures increases. In addition, the method correctly recognizes a facial expression that continuously follows another one. Takahiro Otsuka, Jun Ohya |
ICIP (2) | 2 |
| 1997 | An Affine Coordinate Based Algorithm for Reprojecting the Human Face for Identification TasksabstractWe present an algorithm to generate new views of a human face, starting with at least two other views of the face. In a typical face recognition system, the task of comparison becomes easier if the faces have similar orientation with respect to the camera. The affine coordinate based reprojection algorithm presented enables us to do that. Dense point matches between the two input faces of the same individual are computed using an affine coordinate based reprojection framework. This is followed by the reprojection of one of these to faces to the target face once the user has matched four feature points across two input face images and the target face image. Kuntal Sengupta, Jun Ohya |
ICIP (3) | 2 |
| 1997 | Hand Image Segmentation Using Sequential-Image-Based Hierarchical AdaptationabstractTwo methods to extract a moving target region from a series of images are presented. Pixel value distributions for both the target object and background region are estimated for each pixel with roughly extracted moving regions. Using the distributions, stable target extraction is performed. In the first method, the distributions are approximated with Gaussian distribution functions and the probability of a pixel being associated with the target object is calculated. In the second method, a Markov random field model is applied to perform region segmentation on regularized input images using the estimated pixel value distributions. The texture parameters for the target object region can be calculated from the estimated pixel value distributions. Experimental results obtained by these two methods using hand motion images are presented. Akira Utsumi, Jun Ohya |
ICIP (1) | 2 |
| 1996 | Automatic extraction and tracking of contoursabstractThis paper considers the problem of extracting and tracking complex contours without user interaction. We assume that a complex contour consists of contour segments whose spatial coordinates and intensity gradient vary smoothly in the direction normal to themselves. In our algorithm, digital curves that could correspond to contour segments are extracted by connecting edge pixels using a B-spline based contour segment model. The extracted curves trace the contour segments at the next frame by using the active contour model technique. Experimental results show even occluded contours can be tracked automatically. Koichi Hata, Jun Ohya, Fumio Kishino, Ryohei Nakatsu |
ICPR | 2 |
| 1996 | Detecting facial expressions from face images using a genetic algorithmabstractA new method to detect deformations of facial parts from a face image regardless of changes in the position and orientation of a face using the genetic algorithm is proposed. Facial expression parameters that are used to deform and position a 3D face model are assigned to the genes of an individual in a population. The face model is deformed and positioned according to the gene values of each individual and is observed by a virtual camera, and a face image is synthesized. The fitness which evaluates to what extent the real and synthesized face images are similar to each other is calculated. After this process is repeated for sufficient generations, the parameter estimation is obtained from the genes of the individual with the best fitness. Experimental results demonstrate the effectiveness of the method. Jun Ohya, Fumio Kishino |
ICPR | 1 |
| 1995 | Virtual-Space Teleconferencing - Real-Time Detection and Reproduction of 3D Face and Body Images
Fumio Kishino, Kazuyuki Ebihara, Jun Ohya |
ACCV | 3 |
| 1995 | Virtual Space Teleconferencing: Real-Time Reproduction of 3D Human Images
Jun Ohya, Yasuichi Kitamura, Fumio Kishino, Nobuyoshi Terashima, Haruo Takemura, Hirofumi Ishii |
J. Vis. Commun. Image Represent. | 1 |
| 1994 | Dense, time-varying range data acquisition from stereo pairs of thermal and intensity imagesabstractWe propose a new method for acquiring time-sequential range images that provide dense range data. In our approach, stereo pairs of thermal and intensity images are synchronously acquired and are mutually registered. The thermal images are segmented into isotemperature regions. Contour-based matching is done for the isotemperature regions in the thermal images independently at each time instant, and by temporal correspondence, possible matching pairs of contours are generated. By evaluating the similarities of the pairs, consistent and likely pairs are chosen. To get dense range data, intensity profiles within the isotemperature regions are matched by dynamic programming. Experiments on real scenes, including a sequence showing a moving human being, show promising results.> Jun Ohya, Fumio Kishino |
CVPR | 1 |
| 1994 | Human posture estimation from multiple images using genetic algorithmabstractA new method for estimating human postures at a time instant from multiple images using a genetic algorithm is proposed. The posture parameters to be estimated are assigned to the genes of individuals in the population. For each individual, its fitness evaluates to what extent the multiple human images synthesized by deforming a 3D human model according to the values of the genes are registered to the real multiple human images. Genetic operations such as natural selection, crossover and mutation are performed, and individuals in the next generation are generated. After a certain number of repetitions for these processes, the estimated parameter values are obtained from the individual with the best fitness. Experiments using synthesized human multiple images show promising results. Jun Ohya, Fumio Kishino |
ICPR (1) | 1 |
| 1994 | Recognizing Characters in Scene ImagesabstractAn effective algorithm for character recognition in scene images is studied. Scene images are segmented into regions by an image segmentation method based on adaptive thresholding. Character candidate regions are detected by observing gray-level differences between adjacent regions. To ensure extraction of multisegment characters as well as single-segment characters, character pattern candidates are obtained by associating the detected regions according to their positions and gray levels. A character recognition process selects patterns with high similarities by calculating the similarities between character pattern candidates and the standard patterns in a dictionary and then comparing the similarities to the thresholds. A relaxational approach to determine character patterns updates the similarities by evaluating the interactions between categories of patterns, and finally character patterns and their recognition results are obtained. Highly promising experimental results have been obtained using the method on 100 images involving characters of different sizes and formats under uncontrolled lighting.> Jun Ohya, Akio Shio, Shigeru Akamatsu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | A new method for acquiring time-sequential range images by integrating stereo pairs of thermal and intensity imagesabstractA new method for acquiring time-sequential range images is proposed. Stereo pairs of thermal and intensity images are synchronously acquired and are mutually registered. Stereo thermal images are segmented into isotemperature regions. Contour based matching is done for the isotemperature regions. To supplement sparse range data obtained from the contour matching, dynamic programming matching is performed for either intensity profiles or edges in the stereo intensity images. By corresponding pixel pairs obtained from the matching processes, the 3-D coordinates of the points can be calculated. Experiments with real scenes having moving human beings show promising results.> Jun Ohya, Fumio Kishino |
CVPR | 1 |
| 1993 | Time-varying homotopy and the animation of facial expressions for 3D virtual space teleconferencingabstractA homotopy describes the transformation of one arbitrary curve into another that shares the same endpoints. In this paper, we propose a deformable cylinder model, based on homotopy, in which an arbitrary surface interpolated between two contours via a blending function is transformed into another surface over time. We then show how this homotopic deformation can be applied to the realistic animation of human faces in a virtual space teleconferencing system. Specifically, we show that facial expressions such as wrinkling of the forehead and opening and closing of the mouth can be synthesized and animated in real time through 3D homotopic deformations. Souichi Kajiwara, Hiromi T. Tanaka, Yasuichi Kitamura, Jun Ohya, Fumio Kishino |
VCIP | 4 |
| 1992 | Smoothed local generalized cones: an axial representation of 3D shapesabstractThe recovery of viewpoint-independent descriptions of 3D shapes from two-and-one-half dimensional images. A novel 3D shape representation called the smoothed local generalized cones (SLGCs) is proposed. This representation is suitable for recovery of the axis, because the local constraint that characterizes a data set corresponding to the same axis point, namely, the local generalized cone (LGC), is explicitly defined. The extracted axis can be used as a basis for determining a natural parameterization of the object surface. Using this parameterization, the deformable surface fitting problem results in a linear least-squares problem, so stable volumetric recovery is possible. Recovery experiments involving real 3D range images are reported.> Yoshinobu Sato, Jun Ohya, Kenichiro Ishii |
CVPR | 2 |
| 1992 | Recovery of hierarchical part structure of 3-D shape from range imageabstractA formulation of the part decomposition problem motived by the minimum-description-length (MDL) criteria is presented. Unlike previous MDL approaches which use analytic functions, a general geometric constraint, convexity, is used as a part constraint. Therefore, the method is suitable for complex natural shapes such as human faces. The recovery process consists of a bottom-up grouping process and a subsequent optimization process based on the MDL criteria. The definite causal relations of part structure between different sensitivity levels are used to recover the hierarchy of part structure. Part decomposition experiments involving real 3-D range images are reported.> Yoshinobu Sato, Jun Ohya, Kenichiro Ishii |
CVPR | 2 |
| 1992 | Recognizing human action in time-sequential images using hidden Markov modelabstractA human action recognition method based on a hidden Markov model (HMM) is proposed. It is a feature-based bottom-up approach that is characterized by its learning capability and time-scale invariability. To apply HMMs, one set of time-sequential images is transformed into an image feature vector sequence, and the sequence is converted into a symbol sequence by vector quantization. In learning human action categories, the parameters of the HMMs, one per category, are optimized so as to best describe the training sequences from the category. To recognize an observed sequence, the HMM which best matches the sequence is chosen. Experimental results for real time-sequential images of sports scenes show recognition rates higher than 90%. The recognition rate is improved by increasing the number of people used to generate the training data, indicating the possibility of establishing a person-independent action recognizer.> Junji Yamato, Jun Ohya, Kenichiro Ishii |
CVPR | 2 |
| 1988 | A relaxational extracting method for character recognition in scene imagesabstractAn effective extraction algorithm for character recognition in scene images is developed. Character candidate regions are detected by an image segmentation method based on adaptive thresholding and by evaluating gray-level difference between adjacent regions. Character pattern candidates are obtained by associating the detected regions according to their positions and gray levels and evaluating the aspect ratio of the rectangle that circumscribes the associated regions. The character recognition process selects character pattern candidates with high similarities to any categories in a dictionary (high similarity patterns). Highly promising experimental results have been obtained using the method on 100 images acquired under uncontrolled lighting. This algorithm can be applied to read characters of any size and format under various lighting conditions.> Jun Ohya, Akio Shio, Shigeru Akamatsu |
CVPR | 1 |