Satoshi Nishiguchi

dblp:73/1988 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 A Semi-automatic Quality Assessment System for Capturing High-quality Fundus Image
abstract
The precision of medical diagnoses based on images is inextricably linked to the quality and clarity of the images. Poor image quality can impede accurate diagnosis and pose challenges for both physicians and machine learning algorithms in interpreting images. By automating the process of fundus image quality assessment during capture, we can ensure that only high-quality images are used, improving diagnostic accuracy. Accordingly, we propose a novel approach for automatically assessing the quality of fundus images using deep learning techniques. Our method incorporates retinal vessel segmentation into RGB images to create four-channel images and then trains a deep learning model on these images to identify the image focus and quality of the images. It can evaluate the quality of fundus images and request image recapture with the adjustment of specific parameters that are classified as poor quality. Our proposed method has the potential to improve the diagnostic accuracy and efficiency of retinal disease diagnosis, particularly in telemedicine settings. By automating the process of fundus image quality assessment, we can ensure that only high-quality images are used for diagnosis, thus improving diagnostic precision. It can serve as an efficient screening tool in the initial stage of acquiring high-quality fundus images.
Asif Mohammed Arfi, Masahiro Toyoura, Kenji Kashiwagi, Satoshi Nishiguchi, Kentaro Go, Zhenyang Zhu, Xiaoyang Mao
CW4
2024 Object Pose Estimation for Grasping by Autonomous Mobile Robots
abstract
The objective of this research is to ascertain the posture of an object, which is one of the essential pieces of information required for grasping an object. This will be achieved by capturing images of the object with a single camera mounted on an autonomous mobile robot. As a single camera is unable to observe an object from multiple directions, data from the other side, which is not captured by the camera, cannot be obtained. As a result, it is challenging to obtain a comprehensive understanding of the object’s shape. Therefore, we propose a methodology for fitting a three-dimensional shape model of the target to the point cloud data obtained via a depth camera capable of observing the surrounding three-dimensional shape. This paper demonstrates the efficacy of the proposed method for estimating the attitude of empty cans on the road, which are the target of refuse collection operations. This is achieved by modeling the shape of a typical empty can and measuring the similarity between the captured 3D shape and the model shape.
Yuki Minamida, Satoshi Nishiguchi
CW2
2024 Whole Slide Image Annotation Support for Estimating Lesion Proportions
abstract
The annotation of medical images for the purpose of segmentation demands a high level of expertise and experience, and the generation of suitable datasets represents a significant challenge. In this study, we propose a method that facilitates annotation on a whole slide image (WSI) without requiring detailed specification of the lesion area. The objective is to enable annotation without the need for specialised knowledge. In particular, the WSI is divided into smaller regions for annotation purposes, with the proportion of lesions in each region being compared with reference images of similar regions, for which the proportion of lesions is known. The proportion of lesions in the reference small region image deemed to be most analogous is then designated as the annotation. The outcomes of an annotation experiment on multiple subjects based on this approach indicated that as the level of experience of the annotator increased, the variability in the accuracy rate diminished. This suggests that the annotators had acquired knowledge and skills through the process of annotation.
Nana Takano, Satoshi Nishiguchi, Masahiro Toyoura
CW2
2024 Natural Operation of Victim Avatars in Temporary Housing in the Metaverse
abstract
This paper proposes a method for controlling disaster victim avatars in nursing education materials using the metaverse. The method combines animation with control by a wearable motion capture device to achieve natural movements. While metaverse-based simulators have been developed for various fields, they often lack the naturalness and fidelity of human models. To address this, we developed a prototype system that enhances avatar movements by integrating animation with motion capture, and evaluated the avatars’ impressions.
Soichi Takeuchi, Satoshi Nishiguchi, Yasuharu Mizutani, Wataru Hashimoto 0003, Yukari Kamei, Yumiko Matsushita
CW2
2019 A Study of Usability Improvement in Immersive VR Programming Environment
abstract
Visual programming environments such as Scratch have been proposed for beginners. In those environments, programming is possible by arranging function blocks expressed on two dimensions. In order to improve the browsability of many blocks, an environment for programming by arranging functional blocks with hand gesture interaction in immersive VR space has been proposed. However, operation error by Leap Motion when setting the value and the usability of the environment was not good due to the error of estimation of hand gesture. In this study, we propose an operation method using VR controller for programming in an immersive VR environment and compare the results. We will incorporate these results into future development.
Atsuki Onishi, Satoshi Nishiguchi, Yasuharu Mizutani, Wataru Hashimoto 0003
CW2
2017 Segmentation and tracking of object when grasped and moved within living spaces
abstract
Object tracking from camera images is a key technique in the field of computer vision. Because the most representative application of this technique is security or surveillance, objects such as pedestrians or vehicles in public spaces are considered as tracking targets. However, it is also useful to track objects such as daily items found in indoor living spaces, including offices, private rooms, and other locations, when they go missing or need to be located. However, it is not known in advance which part of the living space can move as a single object to be tracked. Moreover, because objects in a living space do not move by themselves but are moved by human hands, the bodies of objects in motion are usually occluded by the hands for a relatively long period of time, and drastic changes in the position and appearance of the objects may occur after a long-term occlusion. In this article, we propose an approach to finding the correspondence between objects observed before and after their occlusion for object tracking purposes by detecting the grasp and release of such objects by the same hand. Segmentation of the body of each object is also realized through the detection of its grasp and release.
Takuya Omi, Koh Kakusho, Masaaki Iiyama, Satoshi Nishiguchi
SMC4
2017 Estimating the target of interaction for each human in office space with obstacles using 3D observation
abstract
Conventional studies on human behavior recognition have mainly focused on individual actions, including facial expressions and postures. However, most human behavior involves face-to-face interactions with other humans or objects, such as PCs. The previous work on recognizing face-to-face interaction focused on recognition of each group of humans engaged in conversation in an open space, which is considered a free space without obstacles. However, face-to-face conversation in our daily activities also occurs in a space such as an office with various obstacles, which could interfere with the interaction. To cope with this interference, the arrangement of obstacles and humans in the space should be considered. Moreover, in an office space, humans interact not only with other humans but also with certain interactive objects, such as PCs. This study aims to detect the target of interaction of each person in a space with obstacles by considering the 3D positions and orientations of humans, interactive objects, and other obstacles. The 3D arrangement of humans and obstacles is obtained by observing the space with an RGB-D camera.
Masatoshi Tsukamoto, Koh Kakusho, Masaaki Iiyama, Satoshi Nishiguchi
SMC4
2013 ActVis: Activity Visualization in Videos
abstract
We present ActVis, which is a computer-aided video surveillance system for detecting and visualizing the activation levels of multiple objects in a video. ActVis indicates "something is happening" in a video. A user arranges panels indicating the regions of focusing objects on the video screen. Temporal differential as an activation level in a panel is detected by the system, and a corresponding seek bar representing the level is generated. In general, high-level features, such as body posture or facial direction/expression, cannot be extracted when the target object is partially occluded in video, or it is not human. By employing the temporal differential as a low-level feature and the metaphor of a level meter, our system can notify a user "when something happens." The user can explore high-level features of the moment. Potential applications of ActVis include the analysis of student activation levels in classroom for professional development of faculty, and observations of wild animals for ecological investigation.
Masahiro Toyoura, Satoshi Nishiguchi, Xiaoyang Mao, Masayuki Murakami
CW2
2010 Extraction of Mastication in Diet Based on Facial Deformation Pattern Descriptor
abstract
In this paper, we describe a method for extraction of mastication from image sequence. Mastication is the first step of eating and is very important. However, people do not need strong mastication muscle because the advancement of cooking and food processing technology makes soft foods. The weak mastication ability of recent human is about to become a serious problem. It will be a risk factor of many diseases. To prevent this, analysis method of mastication is essential. Several studies have been conducted in the research field, but those were not applicable to eating in daily life, because of the various restrictions. We proposed a mastication analysis method using only monocular camera. The key point is Facial Deformation Pattern Descriptor, FDPD which can represent a pattern of facial deformation. By using the FDPD, we could extract mastication from video successfully, and could develop some useful mastication analysis system for healthy eating life.
Kenzaburo Miyawaki, Satoshi Nishiguchi, Mutsuo Sano
ISM2
2009 Synthesizing a high-resolution image of a lecture room using lecturer tracking camera and planar object capturing camera
abstract
We propose a method for synthesizing an undistorted highresolution image of a lecture room. A certain degree of high-resolution videos can be captured using commercially available high-resolution cameras. However, the resolution of objects captured by these cameras is not enough for viewers of a lecture video. To get sufficient high-resolution lecture videos, we synthesize multiple images captured by low cost cameras using a planar perspective projection. However, lecturers are distorted by the projection because they are not planar objects and do not exist on planar objects in a lecture room. To get an undistorted figure of a lecturer, we introduce a fixed-viewpoint projection using a lecturer tracking camera. As a result, we could get more high-resolution images including an undistorted figure of a lecturer than ones captured by a standard high-resolution camera.
Satoshi Nishiguchi, Yoshitaka Morimura, Takehiro Yashiro, Koh Kakusho, Michihiko Minoh
ICME1
2004 Active Lighting for Object Brightness Control
abstract
We present an active lighting method that can control brightness of multiple objects in a scene to any values on video images. Our approach does not modify pixel values such as in retouching but changes intensities of multiple lights in the scene. Our algorithm needs little computation to calculate appropriate intensities of the lights because we formulate the brightness of the objects so that the essential parameters in our formulation can be measured in advance. We have implemented a prototype system that can control brightness of faces of two people in a room to intended values with eight controllable lights as they are walking.
Yoshinari Kameda, Jun Shingu, Satoshi Nishiguchi, Michihiko Minoh
CW3
2003 CARMUL: concurrent automatic recording for multimedia lecture
abstract
An advanced multimedia lecture recording system is presented in this paper. The purpose of our system is to capture multimodal information that can be received only when lectures are being held in a classroom where a teacher and students share the same time and space. The system captures not only hand-writings and slide switching intervals but also audio and video of the people with their spatial location information. These recorded media will be served to users in distance learning process. We designed the system so that it does not interfere with classes. It can archive lectures while they hold classes on regular basis and generate a multimedia archive automatically. Teachers are neither asked to remain at a certain place nor wired with devices they would have to put on. We have implemented our approach and our lecture archive system currently works for six classes per week, since October, 2002.
Yoshinari Kameda, Satoshi Nishiguchi, Michihiko Minoh
ICME2
2003 A sensor-fusion method for detecting a speaking student
abstract
In this paper, we propose a method for detecting the location of the speaker that is a target of automatic video filming in distance learning and lecture archive. It is required that a face of a speaking student is filmed in a lecture video. For this purpose, it is necessary to detect the location of a speaker. An acoustic sensor such as a microphone array is used widely to detect the location of a sound source. However, it is difficult to detect the location of a sound source precisely using only microphone array because of sound noise in a large space such as a lecture room. In this paper, we propose a method for detecting more precise location of a speaker in the lecture room using not only the microphone array but also visual sensors. The result shows that the precision ratio of detecting the location of a speaker was improved about 20% by our sensor-fusion method.
Satoshi Nishiguchi, Kazuhide Higashi, Yoshinari Kameda, Michihiko Minoh
ICME1
2003 Environmental Media - In the Case of Lecture Archiving System
Michihiko Minoh, Satoshi Nishiguchi
KES2