Daisuke Deguchi

dblp:00/2175 · DBLP profile ↗
← Back
77ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0003-0603-8790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 31 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021
YearPublicationVenuePosition
2026 Capture-Calibrate-Coach: A Graph-Based Framework for Knowledge Monitoring Estimation and Adaptive Feedback
Li Chen 0032, Cheng Tang 0001, Boxuan Ma, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
AIED6
2026 Resolving the Inherent Contextual Insufficiency in Referring Image Segmentation with Global Semantic Priors
Chong Yi, Jialei Chen 0001, Seigo Ito, Hiroshi Murase, Daisuke Deguchi
ICPR (6)5
2026 Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Seigo Ito, Hiroshi Murase
Int. J. Comput. Vis.2
2026 Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi
Int. J. Comput. Vis.7
2026 Correction: Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi
Int. J. Comput. Vis.7
2026 CLIP-to-Seg Distillation for Zero-Shot Semantic Segmentation
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Daisuke Deguchi, Hiroshi Murase
IEEE Trans. Circuits Syst. Video Technol.5
2025 From Reflections to Motifs: A Graph-Based Analysis of Learners' Knowledge Construction
Li Chen 0032, Cheng Tang 0001, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
AIED (6)4
2025 Ranking-Based At-Risk Student Prediction Using Federated Learning and Differential Features
Shunsuke Yoneda, Valdemar Svábenský, Daisuke Deguchi, Atsushi Shimada 0001
EDM4
2025 Single-agent vs. Multi-agent LLM Strategies for Automated Student Reflection Assessment
Li Chen 0032, Cheng Tang 0001, Valdemar Svábenský, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
PAKDD (5)5
2025 Semantic matters: A constrained approach for zero-shot video action recognition
abstract
Zero-shot video action recognition has advanced significantly due to the adaptation of visual-language models, such as CLIP, to video domains. However, existing methods attempt to adapt CLIP to video tasks by leveraging temporal information, neglecting the semantic information (i.e. the latent categories and their relationships) within videos. In this paper, we propose a Semantic Constrained CLIP (SC-CLIP) approach that leverages semantic information to adjust CLIP for video recognition while ensuring its performance on unseen data. SC-CLIP comprises a semantic-related query generation module and a semantic constrained cross attention module. First, the semantic-related query generation module clusters dense tokens from CLIP to generate semantic-related mask. The semantic-related query is then derived by pooling the adapted CLIP output using the semantic-related mask. Next, the semantic constrained cross attention module feeds the generated semantic-related query back into CLIP to probe semantic-related values, enhancing their ability to leverage the vision-language matching capabilities of CLIP. By generating semantic-related query, the semantic information aids in distinguishing similar actions, thereby improving performance on unseen samples. Experimental results on three zero-shot action recognition benchmarks show improvements of up to 1.9% and 2% in harmonic mean under two settings. Code is available at https://github.com/quanzhenzhen/SC-CLIP .
Zhenzhen Quan, Jialei Chen 0001, Daisuke Deguchi, Hiroshi Murase
Pattern Recognit.3
2024 Early Detection of At-risk Students Through Leaning-Activity Forecasting
abstract
With the widespread adoption of digital technologies such as digital textbooks, it has become feasible to collect daily logs of students' learning activities. Accordingly, there has been a growing trend in research using these logs. One of these areas is focusing on predicting grade of each student based on these learning activity logs. However, previous research focused on detecting At-risk students when learning activity logs for all lectures are available. This is not applicable for detection at the first few lectures (i.e. weeks) required in practical usage scenarios. We call this scenario as "early detection" in this paper. However, in early detection, the accuracy of at-risk detection tends to decrease. To solve this problem, we propose a Learning-activity Forecasting Network (LFNet) that improves the accuracy of early detection by aligning the embedding of the first few lectures with that of all lectures. Through experiments on learning activity logs of actual lectures, we confirmed that the proposed method could achieve high At-risk detection accuracy even from the first few lectures of learning activity logs.
Yuya Ozaki, Daisuke Deguchi, Haruya Kyutoku, Hiroshi Murase
ICCE2
2024 LLM-Driven Ontology Learning to Augment Student Performance Analysis in Higher Education
Cheng Tang 0001, Li Chen 0032, Daisuke Deguchi, Takayoshi Yamashita, Atsushi Shimada 0001
KSEM (3)4
2024 Computational measurement of perceived pointiness from pronunciation
abstract
Abstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.8
2024 Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.8
2024 Frozen is better than learning: A new design of prototype-based classifier for semantic segmentation
abstract
Semantic segmentation models comprise an encoder to extract features and a classifier for prediction. However, the learning of the classifier suffers from the ambiguity which is caused by two factors: (1) the weights of a classifier for similar categories may have positive similarities lowing the performance for similar categories, named correlation ambiguity, and (2) the classifier is prone to predict the category with a larger ℓ2 norm and vice versa, termed prior ambiguity. To comedy the issues, we propose Category-Basis Prototype (CBP), frozen and mutually orthogonalized prototypes with equalℓ2 norm. Orthogonalization prevents the prototypes from being similar to each other and the equality decouples the prediction from the ℓ2 norm. To better shape the feature space, we propose Online Centroid Contrastive Loss (OCCL) equipped with centroid and category-level losses. Experiments show that our method yields compelling results over two widely applied benchmarks indicating the effectiveness of our methods.
Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Hiroshi Murase
Pattern Recognit.2
2024 Texture-Guided Transfer Learning for Low-Quality Face Recognition
abstract
Although many advanced works have achieved significant progress for face recognition with deep learning and large-scale face datasets, low-quality face recognition remains a challenging problem in real-word applications, especially for unconstrained surveillance scenes. We propose a texture-guided (TG) transfer learning approach under the knowledge distillation scheme to improve low-quality face recognition performance. Unlike existing methods in which distillation loss is built on forward propagation; e.g., the output logits and intermediate features, in this study, the backward propagation gradient texture is used. More specifically, the gradient texture of low-quality images is forced to be aligned to that of its high-quality counterpart to reduce the feature discrepancy between the high- and low-quality images. Moreover, attention is introduced to derive a soft-attention (SA) version of transfer learning, termed as SA-TG, to focus on informative regions. Experiments on the benchmark low-quality face DB's TinyFace and QMUL-SurFace confirmed the superiority of the proposed method, especially more than 6.6% Rank1 accuracy improvement is achieved on TinyFace.
Meng Zhang 0042, Rujie Liu, Daisuke Deguchi, Hiroshi Murase
IEEE Trans. Image Process.3
2024 Toward Explainable End-to-End Driving Models via Simplified Objectification Constraints
abstract
The end-to-end driving models (E2EDMs) convert environmental information into driving actions using a complex transformation which makes E2EDMs have high prediction accuracy. Due to the black-box nature of transformation, the E2EDMs have low explainability. To solve this problem, explanation methods are used to generate explanations for observation. Based on current explanation methods, previous studies tried to further improve the explainability of E2EDMs by integrating an object detection module, however, these methods have many problems: Firstly, due to the requirement of the object detection module, they lack flexibility. Secondly, they neglect an essential property,i.e., simplicity, to improve explainability. In this paper, since humans prefer object-level and simple explanations in driving tasks, we argue that explainability is decided by two properties which are the objectification degree (the extent to which driving related-object features are utilized) and simplification degree (the simplicity of the explanation), thus we propose Simplified Objectification Branches (SOB) to improve the explainability of E2EDMs. Firstly, this structure could be integrated into any existing E2EDMs and thus have high flexibility. Secondly, the SOB explicitly improves the simplification degree without sacrificing the objectification degree of the explanations. By designing several indicators,i.e., heatmap satisfaction, driving action reproduction score, deception level,etc., we proved that SOB could help E2EDMs generate better explanations. Notably, the SOB could also further enhance E2EDMs’ prediction accuracy.
Daisuke Deguchi, Jialei Chen 0001, Hiroshi Murase
IEEE Trans. Intell. Transp. Syst.2
2023 Refined Objectification for Improving End-to-End Driving Model Explanation Persuasibility*
abstract
With the rapid development of deep learning, many end-to-end autonomous driving models with high prediction accuracy are developed. However, since autonomous driving technology is closely related to human life, users need to be convinced that end-to-end driving models (E2EDMs) not only have high prediction accuracy in known scenarios but also in practice for unknown scenarios. Therefore, engineers and end-users need to grasp the calculation methods of the E2EDMs based on the driving models’ explanations and ensure the explanations are satisfactory. However, few studies have focused on improving the explanation excellence.In this study, among many properties, we aim to improve the persuasibility of the explanation, we propose ROB (refined objectification branches), a structure that could be mounted to any type of existing E2EDMs. By persuasibility evaluation experiments, we demonstrate that one could improve the persuasibility of the explanations by mounting ROB to the E2EDM. As shown in Fig. 1, the focus area of the driving model accurately shrinks to the important and concise objects on account of ROB. We also perform an ablation study to further discuss each branch’s influence on persuasibility. In addition, we test ROB on multiple mainstream backbones and demonstrate that our structure could also improve the model’s prediction accuracy.
Daisuke Deguchi, Hiroshi Murase
IV2
2022 Detection of Birds in a 3D Environment Referring to Audio-Visual Information
abstract
We propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed.
Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi
AVSS7
2021 Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
MMM (2)7
2021 Soft-Boundary Label Relaxation with class placement constraints for semantic segmentation of the railway environment
Yuki Furitsu, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Hiroki Mukojima, Nozomi Nagamine
Pattern Recognit. Lett.2
2020 LFIR2Pose: Pose Estimation from an Extremely Low-resolution FIR image Sequence
abstract
In this paper, we propose a method for human pose estimation from a Low-resolution Far-InfraRed (LFIR) image sequence captured by a 16 × 16 FIR sensor array. Human body estimation from such a single LFIR image is a hard task. For training the estimation model, annotation of the human pose to the images is also a difficult task for human. Thus, we propose the LFIR2Pose model which accepts a sequence of LFIR images and outputs the human pose of the last frame, and also propose an automatic annotation system for the model training. Additionally, considering that the scale of human body motion is largely different among body parts, we also propose a loss function focusing on the difference. Through an experiment, we evaluated the human pose estimation accuracy with an original data set, and confirmed that human pose can be estimated accurately from an LFIR image sequence.
Saki Iwata, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa
ICPR3
2020 Ω-GAN: Object Manifold Embedding GAN for Image Generation by Disentangling Parameters into Pose and Shape Manifolds
abstract
In this paper, we propose Object Manifold Embedding GAN (Ω-GAN) to generate images of variously shaped and arbitrarily posed objects from a noise variable sampled from a distribution defined over the pose and the shape manifolds in a vector space. We introduce Parametric Manifold Sampling to sample noise variables from a distribution over the pose manifold to conditionally generate object images in arbitrary poses by tuning the pose parameter. We also introduce Object Identity Loss for clearly disentangling the pose and shape parameters, which allows us to maintain the shape of the object instance when only the pose parameter is changed. Through evaluation, we confirmed that the proposed Ω-GAN could generate variously shaped object images in arbitrary poses by changing the pose and shape parameters independently. We also introduce an application of the proposed method for object pose estimation, through which we confirmed that the object poses in the generated images are accurate.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICPR2
2020 Median-Shape Representation Learning for Category-Level Object Pose Estimation in Cluttered Environments
abstract
In this paper, we propose an occlusion-robust pose estimation method of an unknown object instance in an object category from a depth image. In a cluttered environment, objects are often occluded mutually. For estimating the pose of an object in such a situation, a method that de-occludes the unobservable area of the object would be effective. However, there are two difficulties; occlusion causes the offset between the center of the actual object and its observable area, and different instances in a category may have different shapes. To cope with these difficulties, we propose a two-stage Encoder-Decoder model to extract features with objects whose centers are aligned to the image center. In the model, we also propose the Median-shape Reconstructor as the second stage to absorb shape variations in a category. By evaluating the method with both a large-scale virtual dataset and a real dataset, we confirmed the proposed method achieves good performance on pose estimation of an occluded object from a depth image.
Hiroki Tatemichi, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Ayako Amma, Hiroshi Murase
ICPR3
2020 Imageability Estimation using Visual and Language Features
abstract
Imageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability.
Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
ICMR7
2020 Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
MMM (2)5
2020 More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
MMM (2)6
2020 Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.6
2019 Exemplar-Based Pseudo-Viewpoint Rotation for White-Cane User Recognition from a 2D Human Pose Sequence
abstract
In recent years, various facilities are equipped to support visually impaired people, but accidents caused by visual disabilities still occur. In this paper, to support the visually-impaired people in a public space, we aim to classify whether a pedestrian image sequence obtained by a surveillance camera is a white-cane user or not from the temporal transition of a human pose represented as 2D coordinates. However, since the appearance of the 2D pose varies largely depending on the viewpoint of the pose, it is difficult to classify them. So, in this paper, we propose a method to rotate the viewpoint of a pose from various pseudo-viewpoints based on a pair of 2D poses simultaneously observed and classify the sequence by multiple classifiers corresponding to each viewpoint. Viewpoint rotation makes it possible to obtain pseudo-poses seen from various pseudo-viewpoints, extract richer pose features, and recognize white-cane users more accurately. Through an experiment, we confirmed that the proposed method improves the recognition rate by 12% compared to the method not employing viewpoint rotation.
Naoki Nishida 0003, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Jun Piao
AVSS3
2019 Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.5
2018 Development of "KamiRepo" system with automatic student identification to handle handwritten assignments on LMS
abstract
Learning management systems (LMSs) have become fundamental tools for higher education, and frameworks to leverage digital education data within an LMS have attracted attention. On the other hand, there is strong demand for dealing with various education data provided not only from electronic media but also from non-electronic media, such as handwritten assignments. In addition, it is desirable to reduce time-consuming tasks such as sorting and returning handwritten assignments by lecturers. With this background, this paper describes the development of "KamiRepo1", a system that makes it possible to automatically upload handwritten assignments to an LMS. In this system, a deep-learning-based retrainable optical character recognition (OCR) system is developed to identify scanned handwritten assignments of individual students and read their scores. Then, their scanned files, automatically separated from the entire file of scanned handwritten assignments, are returned to the individual students through the LMS together with their corresponding scores. Compared with a conventional system using a dedicated multifunction printer, our developed system is capable of 1) using general-purpose scanners, 2) using a user interface on a Web browser, and 3) achieving accurate student identification. We launched this system in our university in April 2017 and have evaluated its effectiveness. The experimental results obtained using real data collected for 6 months showed that our system achieved a 99.7% success rate in the automatic upload process, and it was confirmed that the system can greatly reduce the burden of sorting and returning handwritten assignments.
Shunya Seiya, Ryuya Ito, Kosuke Okamoto, Ukyo Tanikawa, Shigeki Ohira, Daisuke Deguchi, Tomoki Toda
EDUCON6
2018 Gaze-Inspired Learning for Estimating the Attractiveness of a Food Photo
abstract
The number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them.
Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM6
2018 Voting-based Hand-Waving Gesture Spotting from a Low-Resolution Far-Infrared Image Sequence
abstract
We propose a temporal spotting method of a hand gesture from a low-resolution far-infrared image sequence captured by a far-infrared sensor array. The sensor array captures the spatial distribution of far-infrared intensity as a thermal image by detecting far-infrared waves emitted from heat sources. It is difficult to spot a hand gesture from a sequence of thermal images captured by the sensor due to its low-resolution, heavy noise, and varying duration of the gesture. Therefore, we introduce a voting-based approach to spot the gesture with template matching-based gesture recognition. We confirm the effectiveness of the proposed temporal spotting method in several settings.
Yasutomo Kawanishi, Chisato Toriyama, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa, Masato Kawade
VCIP4
2017 Action recognition from extremely low-resolution thermal image sequence
abstract
This paper proposes a Deep Learning-based action recognition method from an extremely low-resolution thermal image sequence. The method recognizes daily actions by humans (e.g. walking, sitting down, standing up, etc.) and abnormal actions (e.g. falling down) without privacy concerns. While privacy concerns can be ignored, it is difficult to compute feature points and to obtain a clear edge of the human body from an extremely low-resolution thermal image. To address these problems, this paper proposes a Deep Learning-based action recognition method that combines convolution layers and an LSTM layer for learning spatio-temporal representation, whose inputs are the thermal images and their frame differences cropped by the gravity center of human regions. The effectiveness of the proposed method was confirmed through experiments.
Takayuki Kawashima, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Daisuke Deguchi, Tomoyoshi Aizawa, Masato Kawade
AVSS5
2017 Automatic Selection of Web Contents Towards Automatic Authoring of a Video Biography
abstract
In this paper, we propose a method for image selection using Web image search for automatic video biography authoring. In the proposed method, images are selected from the image search results considering their visual contents for inclusion in the video biography. Through evaluation, we confirmed the effectiveness of the proposed image selection method compared to a baseline method which simply selects the top 1 search result.
Ichiro Ide, Yasutomo Kawanishi, Kyoka Kunishiro, Frank Nack, Daisuke Deguchi, Hiroshi Murase
ISM5
2017 Summarization of News Videos Considering the Consistency of Auditory and Visual Contents
abstract
Since news videos are valuable sources of multimedia information on real-world events, there is a demand for viewing them efficiently. However, there is a problem that summarization methods based on auditory contents do not take into account the visual contents. In the case of news videos, due to its presentation style where audio contents and visual contents do not necessarily come from the same source, this could severely decrease the amount of informative visual contents included in the generated summarized video. Thus, we propose a method for summarizing a sequence of news videos considering the consistency of both auditory and visual contents. The proposed method first selects key-sentences from the auditory contents (Closed Caption) of each news story in the sequence, and then selects a shot within the news story whose "Visual Concepts" detected from the visual contents are the most consistent with the key-phrase. Finally, the audio segment corresponding to each key-phrase is overlapped onto the selected shot, and then concatenated to generate a summarized video. The effectiveness of the proposed method was confirmed on several news topics through a subjective experiment.
Ichiro Ide, Ryunosuke Tanishige, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
ISM6
2017 Monocular localization within sparse voxel maps
abstract
We introduce a method that uses a single camera to localize a vehicle within a pre-constructed map consisting of a voxel occupancy grid and road-line marker positions. Sophisticated mapping hardware is capable of creating high-accuracy 3D maps of road environments, but localizing a vehicle within such maps is one of the challenges at the forefront of automated driving. A solution which is robust to dynamic environments, while using only inexpensive sensors, is a difficult problem. In addition, maps that enable precise localization consume a lot of data which is impractical for the expansive environments encountered in real-world road networks. We show how using the area of edge regions shared between rendered views of a compact voxel map and in-vehicle camera images can be coupled with non-linear optimization methods to determine the camera position and pose.
David Wong 0002, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium3
2017 Proposal of a spectral random dots marker using local feature for posture estimation
abstract
We propose a novel marker for robot's grasping task which has the following three aspects: (i) it is easy-to-find in a cluttered background, (ii) it is calculable for its posture (iii) its size is compact. The proposed marker is composed of a random dots pattern, and uses keypoint detection and a scale estimation by Spectral SIFT for dots detection and data decoding. The data is encoded by the scale size of dots, and the same dots in the marker work for both marker detection and data decoding. As a result, the proposed marker size can be compact. We confirmed the effectiveness of the proposed marker through experiments.
Norimasa Kobori, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
VR2
2017 Regression of feature scale tracklets for decimeter visual localization
David Wong 0002, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase
Image Vis. Comput.2
2016 A classification method of cooking operations based on eye movement patterns
abstract
We are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables.
Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ETRA6
2016 Moving camera background-subtraction for obstacle detection on railway tracks
abstract
We propose a method for detecting obstacles by comparing input and reference train frontal view camera images. In the field of obstacle detection, most methods employ a machine learning approach, so they can only detect pre-trained classes, such as pedestrian, bicycle, etc. This means that obstacles of unknown classes cannot be detected. To overcome this problem, we propose a background subtraction method that can be applied to moving cameras. First, the proposed method computes frame-by-frame correspondences between the current and the reference (database) image sequences. Then, obstacles are detected by applying image subtraction to corresponding frames. To confirm the effectiveness of the proposed method, we conducted an experiment using several image sequences captured on an experimental track. Its results showed that the proposed method could detect various obstacles accurately and effectively.
Hiroki Mukojima, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Masato Ukai, Nozomi Nagamine, Ryuta Nakasone
ICIP2
2016 Misclassification tolerable learning for robust pedestrian orientation classification
abstract
In this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
ICPR2
2016 Parts Selective DPM for detection of pedestrians possessing an umbrella
abstract
In recent years, pedestrian detection from an in-vehicle camera has been attracting attention. However, in the case of a raining situation, the detection accuracy decreases because the head of a pedestrian tends to be occluded by an umbrella. In oder to handle such cases, in this paper, as a variation of the Deformable Part Model (DPM) which is widely used in the field of object recognition, we propose “Parts Selective DPM (PS-DPM)” which selectively chooses the original part filters and additional part filters trained independently. In the detection of pedestrians possessing an umbrella, the selection of head and umbrella parts will make pedestrian detection more robust to the occlusion. We conducted experiments to evaluate the performance of the proposed method. As a result, pedestrian detection with the proposed PS-DPM achieved high detection accuracy in rainy weather, compared with the detection by the conventional DPM. Moreover, we confirmed that it did not decrease the pedestrian detection accuracy in fine weather.
Yuto Shimbo, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium3
2015 Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF camera
abstract
This paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification.
Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
Intelligent Vehicles Symposium3
2014 Spatial People Density Estimation from Multiple Viewpoints by Memory Based Regression
abstract
Crowd analysis using cameras has attracted much attention for public safety and marketing. Among techniques of the crowd analysis, we focus on spatial people density estimation which estimates the number of people for each small area in a floor region. However, spatial people density cannot be estimated accurately for an area far from the camera because of the occlusion by people in a closer area. Therefore, we propose a method using a memory based regression method with images captured from cameras from multiple viewpoints. This method is realized by looking up a table that consists of correspondences between people density maps and crowd appearances. Since the crowd appearances include situations where various occlusions occur, an estimation robust to occlusion should be realized. In an experiment, we examined the effectiveness of the proposed method.
Yoshimune Tabuchi, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takayuki Kurozumi, Kunio Kashino
ICPR3
2014 Scene Duplicate Detection from News Videos Using Image-Audio Matching Focusing on Human Faces
abstract
As one tool for structuring a massive volume of archived news videos based on their semantic contents, this paper proposes a method to detect scene duplicates from news videos. A scene duplicate is a pair of video segments taken at the same event from different viewpoints. Referring to the audio channel is effective to detect scene duplicates regardless of viewpoints, but it cannot be relied on when external audio sources (e.g. Narrations, sound effects) overlap the original one. In contrast, the image channel can be useful in most cases, although significant difference in viewpoints affect the detection. The proposed method integrates the information from these two channels in order to improve the accuracy of scene duplicate detection from news videos. The performance of the proposed method was evaluated through an experiment with actual broadcast news videos. As a result, we obtained the higher detection accuracies in both recall and precision. Therefore, we confirmed the effectiveness of the proposed method.
Haruka Kumagai, Keisuke Doman, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM4
2014 Estimation of traffic sign visibility considering local and global features in a driving environment
abstract
This paper proposes a camera-based visibility estimation method for a traffic sign. The visibility here indicates how a visual target is easy to be detected and recognized by a human driver (not a machine). This research aims at realizing a nuisance-free driver assistance system which sorts out information depending on the visibility of a visual target, in order to prevent driver distraction. Our previous study on estimating the visibility of a traffic sign considered only the effect of the local region around a target, assuming the situation that a driver's gaze is around it. The proposed method integrates both the local features and global features in a driving environment without such an assumption. The global features evaluate the positional relationships between traffic signs and the appearance around the fixation point of a driver's gaze, which considers the effect of the driver's entire field of view. Experimental results showed the effectiveness of incorporating the global features for estimating the visibility of a traffic sign.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Utsushi Sakai
Intelligent Vehicles Symposium2
2014 Single camera vehicle localization using SURF scale and dynamic time warping
abstract
Vehicle ego-localization is an essential process for many driver assistance and autonomous driving systems. The traditional solution of GPS localization is often unreliable in urban environments where tall buildings can cause shadowing of the satellite signal and multipath propagation. Typical visual feature based localization methods rely on calculation of the fundamental matrix which can be unstable when the baseline is small. In this paper we propose a novel method which uses the scale of matched SURF image features and Dynamic Time Warping to perform stable localization. By comparing SURF feature scales between input images and a pre-constructed database, stable localization is achieved without the need to calculate the fundamental matrix. In addition, 3D information is added to the database feature points in order to perform lateral localization, and therefore lane recognition. From experimental data captured from real traffic environments, we show how the proposed system can provide high localization accuracy relative to an image database, and can also perform lateral localization to recognize the vehicle's current lane.
David Wong 0002, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium2
2014 Estimation of the Representative Story Transition in a Chronological Semantic Structure of News Topics
abstract
It is important to track the flow of topics to thoroughly understand the contents. Accordingly, a method that structures the chronological semantic relations between news stories, namely a "topic thread structure" has been proposed. It allows the comprehensive understanding of a topic by chronologically tracking stories one by one from the initial story. However, this task imposes a user to watch many stories when it contains various sub-topics. Thus, we propose a method that estimates the representative story transition in a topic thread structure. In the proposed method, features obtained from a story and those from the topic thread structure are used for the estimation. We confirmed the effectiveness of the proposed method by comparing the results obtained from the proposed method to the ground truth obtained from votes in a subjective experiment.
Kosuke Kato, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ICMR3
2014 Event Detection based on Twitter Enthusiasm Degree for Generating a Sports Highlight Video
abstract
This paper presents a Twitter-based event detection method based on "Twitter Enthusiasm Degrees (TED)" toward generating a highlight video of a sports game. Existing methods not only depend on both languages and sports types but also often falsely detect non-target events. In contrast, the proposed method detects sports events using TEDs calculated from several kinds of string features independent of languages and sports. We applied the proposed method to actual sports games, and compared the detected events with the events present in broadcasted highlight videos, and confirmed the effectiveness and the language and sports type independencies of the proposed method.
Keisuke Doman, Taishi Tomita, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ACM Multimedia4
2013 Pedestrian detection by scene dependent classifiers with generative learning
abstract
Recently, pedestrian detection from in-vehicle camera images is becoming an crucial technology for Intelligent Transportation Systems (ITS). However, it is difficult to detect pedestrians accurately in various scenes by obtaining training samples. To tackle this problem, we propose a method to construct scene dependent classifiers to improve the accuracy of pedestrian detection. The proposed method selects an appropriate classifier based on the scene information that is a category of appearance associated with location information. To construct scene dependent classifiers, the proposed method introduces generative learning for synthesizing scene dependent training samples. Experimental results showed that the detection accuracy of the proposed method outperformed the comparative method, and we confirmed that scene dependent classifiers improved the accuracy of pedestrian detection.
Hidefumi Yoshida, Daichi Suzuo, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takashi Machida, Yoshiko Kojima
Intelligent Vehicles Symposium3
2013 Detection of Biased Broadcast Sports Video Highlights by Attribute-Based Tweets Analysis
Takashi Kobayashi 0001, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
MMM (2)3
2012 Robust Face Super-Resolution Using Free-Form Deformations for Low-Quality Surveillance Video
abstract
Recently, the demand for face recognition to identify persons from surveillance video cameras has rapidly increased. Since surveillance cameras are usually placed at positions far from a person's face, the quality of face images captured by the cameras tends to be low. This degrades the recognition accuracy. Therefore, aiming to improve the accuracy of the low-resolution-face recognition, we propose a video-based super-resolution method. The proposed method can generate a high-resolution face image from low-resolution video frames including non-rigid deformations caused by changes of face poses and expressions without using any positional information of facial feature points. Most existing techniques use the facial feature points for image alignment between the video frames. However, it is difficult to obtain the accurate positions of the feature points from low-resolution face images. To achieve the alignment, the proposed method uses a free-form deformation method that flexibly aligns each local region between the images. This enables super-resolution of face images from low-resolution videos. Experimental results demonstrated that the proposed method improved the performance of super-resolution for actual videos in terms of both image quality and face recognition accuracy.
Tomonari Yoshida, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICME3
2012 Estimation of the human performance for pedestrian detectability based on visual search and motion features
Masashi Wakayama, Daisuke Deguchi, Keisuke Doman, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
ICPR2
2012 Development and comparison of new hybrid motion tracking for bronchoscopic navigation
Xióngbiao Luó, Marco Feuerstein, Daisuke Deguchi, Takayuki Kitasaka, Hirotsugu Takabatake, Kensaku Mori
Medical Image Anal.3
2011 Low Resolution QR-Code Recognition by Applying Super-Resolution Using the Property of QR-Codes
abstract
This paper proposes a method for low resolution QR-code recognition. A QR-code is a two-dimensional binary symbol that can embed various information such as characters and numbers. To recognize a QR-code correctly and stably, the resolution of an input image should be high. In practice, however, recognition of a QR-code is usually difficult due to low resolution when it is captured from a distance. In this paper, we propose a method to improve the performance of low resolution QR-code recognition by using the super-resolution technique that generates a high resolution image from multiple low-resolution images. Although a QR-code is a binary pattern, it is observed as a grayscale image due to the degradation through the capturing process. Especially the pixels around the borders between white and black regions become ambiguous. To overcome this problem, the proposed method introduces a binary pattern constraint to generate super-resolved images appropriate for recognition. Experimental results showed that a recognition rate of 98% can be achieved by the proposed method, which is a 15.7% improvement in comparison with a method using a conventional super-resolution method.
Yuji Kato, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR2
2011 Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News Videos
abstract
We propose a method to detect the inconsistency between a subject and the speaker for extracting speech scenes from news videos. Speech scenes in news videos contain a wealth of multimedia information, and are valuable as archived material. In order to extract speech scenes from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such approach, since news videos contain non-speech scenes where the speaker is not the subject, such as narrated scenes. To solve this problem, we propose a method to discriminate between speech scenes and narrated scenes based on the co-occurrence between a subject's lip motion and the speaker's voice. The proposed method uses lip shape and degree of lip opening as visual features representing a subject's lip motion, and uses voice volume and phoneme as audio feature representing a speaker's voice. Then, the proposed method discriminates between speech scenes and narrated scenes based on the correlations of these features. We report the results of experiments on videos captured in a laboratory condition and also on actual broadcast news videos. Their results showed the effectiveness of our method and the feasibility of our research goal.
Shogo Kumagai, Keisuke Doman, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM4
2011 Intelligent traffic sign detector: Adaptive learning based on online gathering of training samples
abstract
This paper proposes an intelligent traffic sign detector using adaptive learning based on online gathering of training samples from in-vehicle camera image sequences. To detect traffic signs accurately from in-vehicle camera images, various training samples of traffic signs are needed. In addition, to reduce false alarms, various background images should also be prepared before constructing the detector. However, since their appearances vary widely, it is difficult to obtain them exhaustively by manual intervention. Therefore, the proposed method simultaneously obtains both traffic sign images and background images from in-vehicle camera images. Especially, to reduce false alarms, the proposed method gathers background images that were easily mis-detected by a previously constructed traffic sign detector, and re-trains the detector by using them as negative samples. By using retrospectively tracked traffic sign images and background images as positive and negative training samples, respectively, the proposed method constructs a highly accurate traffic sign detector automatically. Experimental results showed the effectiveness of the proposed method.
Daisuke Deguchi, Daisuke Shirasuna, Keisuke Doman, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium1
2011 Estimation of traffic sign visibility considering temporal environmental changes for smart driver assistance
abstract
We propose a visibility estimation method for traffic signs considering temporal environmental changes, as a part of work for the realization of nuisance-free driver assistance systems. Recently, the number of driver assistance systems in a vehicle is increasing. Accordingly, it is becoming important to sort out appropriate information provided from them, because providing too much information may cause driver distraction. To solve such a problem, we focus on a visibility estimation method for controlling the information according to the visibility of a traffic sign. The proposed method sequentially captures a traffic sign by an in-vehicle camera, and estimates its accumulative visibility by integrating a series of instantaneous visibility. By this way, even if the environmental conditions may change temporally and complicatedly, we can still accurately estimate the visibility that the driver perceives in an actual traffic scene. We also investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium2
2011 Road image update using in-vehicle camera images and aerial image
abstract
Road image is becoming important for several applications such as car navigation systems, traffic environment research, city modeling. Usually, a road image can be obtained from an aerial image but the resolution of the aerial image is often low, or it contains occlusions by obstacles. Therefore, the update of road image is required. In this paper, we propose a road image mosaicing method using in-vehicle camera images and an aerial image. We first perform image registration of road regions between these images, and then, we generate a large road image by performing image mosaicing of road regions in invehicle camera images. In an experiment, we achieved resolution improvement and occlusions removal, and also succeeded in update of a large road image.
Masafumi Noda, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Yoshiko Kojima, Takashi Naito
Intelligent Vehicles Symposium3
2011 3-D line segment reconstruction using an in-vehicle camera for free space detection
abstract
Free space detection is very important for vehicle navigation and safe driving. 3-D line segment reconstruction of a street is important for the free space detection because a street-view includes many line segments. For the free space detection, we propose a method for reconstructing 3-D line segments in a streetscape using a monocular in-vehicle camera. The 3-D reconstruction of the line segments is achieved by using each three images from an image sequence. Once accurate camera poses of these images are obtained, one of the remaining crucial problems is to match the line segments between the images correctly. A strategy for finding correspondence of the line segments is as follows: First, the correspondences of line segment candidates are searched by using a two-view constraint. However, the two-view constraint has difficulty on determining an unique correspondence geometrically. Therefore, the candidates of the line segment correspondences are reduced using a three-view constraint. In order to improve the accuracy, the proposed method exploits a color feature of the line segment and a preliminary knowledge of the vehicle motion. Finally, the line segments are reconstructed using the correspondences. From an experimental result, we confirmed the effectiveness of the proposed method. Application to the free space detection demonstrated the usefulness of the reconstructed line segments.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium2
2011 Scene segmentation of wedding party videos by scenario-based matching with example videos
abstract
We propose a method for scene segmentation of a wedding party video. Recently, it has become popular to take videos of a wedding ceremony and its party. Especially, because of its length, each scene of a wedding party video needs to be indexed with each event for efficient browsing. The proposed method segments a wedding party video into scenes of events by scenario- based matching with example videos that are synthesized by combining scenes from other wedding party videos according to a scenario.
Kazuki Sawai, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ACM Multimedia3
2010 Classification of Near-Duplicate Video Segments Based on Their Appearance Patterns
abstract
We propose a method that analyzes the structure of a large volume of general broadcast video data by the appearance patterns of near-duplicate video segments. We define six classification rules based on the appearance patterns of near-duplicate video segments according to their roles, and evaluated them over more than 1,000 hours of actual broadcast video data.
Ichiro Ide, Yuji Shamoto, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ICPR3
2010 Removal of Moving Objects from a Street-View Image by Fusing Multiple Image Sequences
abstract
We propose a method to remove moving objects from an in-vehicle camera image sequence by fusing multiple image sequences. Driver assistance systems and services such as Google Street View require images containing no moving object. The proposed scheme consists of three parts: (i) collection of many image sequences along the same route by using vehicles equipped with an omni-directional camera, (ii) temporal and spatial registration of image sequences, and (iii) mosaicing partial images containing no moving object. Experimental results show that 97.3% of the moving object area could be removed by the proposed method.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICPR2
2010 Multimedia Supplementation to a Cooking Recipe Text for Facilitating Its Understanding to Inexperienced Users
abstract
Assisting culinary activities for inexperienced users has been considered as an important task in most existing works in the field. On the other hand, recipe texts are becoming available on the Internet in increasing numbers. However, they tend to be written simply by mostly non-professional people, and thus are sometimes difficult for an inexperienced person to follow the steps and manage to cook as they are supposed to. In this paper, we propose a method that detects difficult descriptions for an inexperienced user in an existing text recipe, and supplements them with multimedia contents including text information extracted from a large number of recipes, and also images and video clips on certain kinds of cooking operations, to facilitate the understanding of the recipe. Experimental results showed promising ability of the proposed method to assist inexperienced users understand the descriptions in a recipe.
Ichiro Ide, Yuka Shidochi, Yuichi Nakamura 0001, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ISM4
2010 Estimation of traffic sign visibility toward smart driver assistance
abstract
We propose a visibility estimation method for traffic signs as part of work for realization of nuisance-free driving safety support systems. Recently, the number of driving safety support systems in a car has been increasing. As a result, it is becoming important to select appropriate information from them for safe and comfortable driving because too much information may cause driver distraction and may increase the risk of a traffic accident. One of the approaches to avoid such a problem is to alert the driver only with information which could easily be missed. Therefore, to realize such a system, we focus on estimating the visibility of traffic signs. The proposed method is a model-based method that estimates the visibility of traffic signs focusing on the difference of image features between a traffic sign and its surrounding region. In this paper, we investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium2
2009 Low-Resolution Character Recognition by Video-Based Super-Resolution
abstract
In this paper, we propose a method for recognizing low-resolution characters using a super-resolution technique. Although portable digital cameras can be used for camera based character recognition, the captured images contain several types of noises which make the recognition task difficult. We introduce a phase of super-resolution before the recognition to enhance the resolution of images obtained from a video. The proposed method uses the subspace method for the recognition of characters which are integrated from multiple low-resolution characters by the super-resolution technique. Experimental results show that the proposed method improves the recognition accuracy; we confirmed that the recognition rate for the input size of 7 times 7 pixels was 90.35%, and for the input size of 9 times 9 pixels was 99.97%.
Ataru Ohkura, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR2
2009 Labeling News Topic Threads with Wikipedia Entries
abstract
Wikipedia is a famous online encyclopedia. However most Wikipedia entries are mainly explained by text, so it will be very informative to enhance the contents with multimedia information such as videos. Thus we are working on a method to extend information of Wikipedia entries by means of broadcast videos which explain the entries. In this work, we focus especially on news videos and Wikipedia entries about news events. In order to extend information of Wikipedia entries, it is necessary to link news videos and Wikipedia entries. So the main issue will be on a method that labels news videos with Wikipedia entries automatically. In this way, explanations could be more detailed with news videos can be exhibited, and the context of the news events should become easier to understand. Through experiments, news videos were accurately labeled with Wikipedia entries with a precision of 86% and a recall of 79%.
Tomoki Okuoka, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM3
2009 Automated Anatomical Labeling of Bronchial Branches Extracted from CT Datasets Based on Machine Learning and Combination Optimization and Its Application to Bronchoscope Guidance
Kensaku Mori, Shunsuke Ota, Daisuke Deguchi, Takayuki Kitasaka, Yasuhito Suenaga, Shingo Iwano, Yosihnori Hasegawa, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori
MICCAI (1)3
2009 Selective image similarity measure for bronchoscope tracking based on image registration
Daisuke Deguchi, Kensaku Mori, Marco Feuerstein, Takayuki Kitasaka, Calvin R. Maurer Jr., Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori
Medical Image Anal.1
2007 Bronchoscope Tracking Without Fiducial Markers Using Ultra-tiny Electromagnetic Tracking System and Its Evaluation in Different Environments
Kensaku Mori, Daisuke Deguchi, Kazuyoshi Ishitani, Takayuki Kitasaka, Yasuhito Suenaga, Yosihnori Hasegawa, Kazuyoshi Imaizumi, Hirotsugu Takabatake
MICCAI (2)2
2006 Bronchoscope Tracking Based on Image Registration Using Multiple Initial Starting Points Estimated by Motion Prediction
Kensaku Mori, Daisuke Deguchi, Takayuki Kitasaka, Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori, Calvin R. Maurer Jr.
MICCAI (2)2
2005 Hybrid Bronchoscope Tracking Using a Magnetic Tracking Sensor and Image Registration
Kensaku Mori, Daisuke Deguchi, Kenta Akiyama, Takayuki Kitasaka, Calvin R. Maurer Jr., Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori
MICCAI (2)2
2004 Fast and Accurate Bronchoscope Tracking Using Image Registration and Motion Prediction
Jiro Nagao, Kensaku Mori, Tsutomu Enjouji, Daisuke Deguchi, Takayuki Kitasaka, Yasuhito Suenaga, Jun-ichiro Toriwaki, Hirotsugu Takabatake, Hiroshi Natori
MICCAI (2)4
2003 New Image Similarity Measure for Bronchoscope Tracking Based on Image Registration
Daisuke Deguchi, Kensaku Mori, Yasuhito Suenaga, Jun-ichiro Toriwaki, Hirotsugu Takabatake, Hiroshi Natori
MICCAI (1)1
2002 Tracking of a bronchoscope using epipolar geometry analysis and intensity-based image registration of real and virtual endoscopic images
Kensaku Mori, Daisuke Deguchi, Jun Sugiyama, Yasuhito Suenaga, Jun-ichiro Toriwaki, Calvin R. Maurer Jr., Hirotsugu Takabatake, Hiroshi Natori
Medical Image Anal.2
2001 A Method for Tracking the Camera Motion of Real Endoscope by Epipolar Geometry Analysis and Virtual Endoscopy System
Kensaku Mori, Daisuke Deguchi, Yasuhito Suenaga, Jun-ichiro Toriwaki, Hirotsugu Takabatake, Hiroshi Natori
MICCAI2