VLDB 2026 Research / reviewers in the wild / expert
Yasutomo Kawanishi
dblp:99/8066
· DBLP profile ↗
61ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0002-3799-4550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REACH: Hand Pose Estimation from Room Corners
Shu Nakamura, Ryo Kawahara, Genki Kinoshita, Ryosuke Hirai, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
FG | 5 |
| 2026 | Lightweight Gating Mechanism for RNNs with Sech-Based Vector Gates
Tomohiro Fujita, Yasutomo Kawanishi |
ICPR (12) | 2 |
| 2026 | Frame-Level Driver Drowsiness Detection with Deep Embedded Clustering-based Pseudo Label Refinement
Vijay John, Yasutomo Kawanishi |
IV | 2 |
| 2026 | View-aware Cross-modal Distillation for Multi-view Action RecognitionabstractThe widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen–Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
WACV | 2 |
| 2026 | Hierarchical graph attention networks with spatio-temporal class tokens for distributed audio-visual event classificationabstractAbstract This paper presents a novel multi-view multimodal graph learning framework for distributed audio-visual event classification using synchronized sequences from multi-microphone and multi-camera sensors. Existing approaches often rely on simple aggregation strategies for multi-view multimodal inputs, which fail to adequately capture the complex spatio-temporal relationships both within and across modalities. To address this limitation, we propose a graph attention network architecture with individual frame-level sensor nodes for each microphone and camera, and three types of frame-level spatio-temporal nodes. In this framework, within each temporal frame, audio spatio-temporal nodes connect to microphones, video spatio-temporal nodes to cameras, and audio-video spatio-temporal nodes to all sensor nodes. Temporal edges further interconnect each spatio-temporal node with its corresponding node in preceding frames. This architecture enables dynamic aggregation of sensor node features through attention-weighted mechanisms, generating updated spatio-temporal nodes that capture intra-modal, inter-modal, intra-frame, and inter-frame relational dependencies. Experimental results on the MM-Office and MM-OR datasets demonstrate that the proposed framework significantly outperforms existing baseline methods for audio-visual event classification, highlighting its superior capability in modeling complex spatio-temporal dependencies across distributed sensor networks. Vijay John, Yasutomo Kawanishi |
Multim. Tools Appl. | 2 |
| 2026 | MultiSensor-Home: Multi-modal multi-view dataset and benchmarks for action recognition in home environmentsabstractMulti-modal multi-view action recognition is a rapidly growing area in computer vision, with important applications in surveillance, smart homes, and assistive robotics. However, existing datasets often fail to capture real-world challenges such as distributed sensor layouts, asynchronous data streams, and limited frame-level annotations. To address these limitations, we introduce MultiSensor-Home, a novel multi-modal multi-view dataset specifically designed for realistic residential environments. It comprises 5,250 untrimmed videos recorded in two distinct residential environments, Home-1 and Home-2, using five distributed RGB and audio sensor units, where each video contains multiple sequential actions along with background segments. Each frame in the recorded sequences is manually annotated with fine-grained frame-level action labels, making it, to the best of our knowledge, the first densely annotated multi-view dataset for home activity recognition. To benchmark this dataset, we further propose Act ion Selection Learning-guided Transformer-based Sensor Fusion (ActFusion), a unified method that jointly models temporal dynamics and cross-view correspondence. It dynamically models cross-view relationships and selects informative frames, enabling robust training under both frame-level supervision, where the start and end timings of each action are labeled, and sequence-level supervision, where only action labels are provided. To support reproducible evaluation, we establish a comprehensive benchmark with standardized training and testing protocols. Extensive experiments on MultiSensor-Home and the existing MM-Office datasets show that ActFusion consistently outperforms baseline methods across diverse scenarios. By capturing the challenges in a realistic setting, MultiSensor-Home sets a new benchmark and encourages future research on robust and generalizable action recognition methods. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
Pattern Recognit. | 2 |
| 2026 | Hierarchical Global-Local Fusion for One-stage Open-vocabulary Temporal Action DetectionabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) extends the detection scope of Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) to unseen action classes specified by vocabularies not included in the training data, within untrimmed video. Typical Open-vocab TAD methods adopt a two-stage approach that first proposes candidate action intervals and then identifies those actions. However, errors in the first stage can affect the subsequent stage and the final detection results. Moreover, conventional methods for temporal context analyses tend to focus solely on either global or local context. Focusing solely on the global context can lead to lack of momentary detail, making it difficult to distinguish one action from another. Conversely, focusing only on the local context makes it challenging to determine the start and end timings of action intervals. To address these challenges, we introduce a one-stage approach named Hierarchical Open-vocab TAD (HOTAD), consisting of two branches: Temporal Context Analysis (TCA) and Video–Text Alignment (VTA). The former utilizes Hierarchical Encoder (HE) to fuse global and local temporal features, enabling a comprehensive capture of temporal actions, while the latter branch exploits the synergy between visual and textual modalities for precisely detecting unseen actions in the Open-vocab setting. Experiments and in-depth analysis using the widely recognized datasets THUMOS14 and ActivityNet-1.3 are performed to show the effectiveness of HOTAD. The results highlight remarkable accuracy in detecting a wide range of unseen actions. Furthermore, HOTAD significantly reduces wrong labels and localizes action instances with high precision, showcasing its robustness in complex and dynamic video settings. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Modelling Spatio-Temporal Dynamics by Graph Attention Network for Distributed Multi-Microphone Sound Event ClassificationabstractThis paper introduces a novel distributed multi-microphone sound event classification framework that uses graph attention networks to model spatial and temporal relationships between distributed multi-microphones. Existing methods utilize naive aggregation approaches like concatenation, averaging, or maximum operations for multiple sensor inputs resulting in information loss and limitations in capturing the complex spatial and temporal relationships in the multi-microphone sequence. To address this, we propose a framework based on the graph attention network to model the complex multi-microphone relationship. Our framework is based on two graph structures: a spatio-temporal graph (STG), which is algorithmically modeled to capture inter-microphone and inter-frame relationships, and a learnable fully-connected spatial graph (FCSG), which is designed to capture complementary details. Utilizing them, a graph attention network-based aggregation module effectively updates the graph nodes resulting in improved event classification accuracy. Experimental results on the MM-Office dataset demonstrate that our proposed framework significantly outperforms baseline methods for the event classification task. Vijay John, Yasutomo Kawanishi |
AVSS | 2 |
| 2025 | Head Orientation Estimation from a Low-Resolution Far-Infrared Image SequenceabstractWe propose a novel method for estimating head orientation from a low-resolution far-infrared (LFIR) image sequence captured by a 32 × 32 FIR sensor array. Estimating head orientation from LFIR images is challenging due to their inherent noise and limited resolution. Manually an-notating head orientation is also difficult, and no suitable dataset currently exists. We propose the LFIR2Head model, which leverages a multimodal-temporal framework for robust estimation. We also introduce the LFIR2Head dataset, which was created using an automatic annotation system. Extensive experiments demonstrate the accuracy of head orientation estimation with the proposed dataset. We con-firmed that head orientation can be estimated accurately from an LFIR image sequence. Taiyo Tamaki, Motoharu Sonogashira, Kazuya Kitano, Yuki Fujimura, Takuya Funatomi, Yasuhiro Mukaigawa, Yasutomo Kawanishi |
AVSS | 7 |
| 2025 | Cross-modal Emotion-specific Attention model for Multimodal Emotion RecognitionabstractEmotion recognition plays a crucial role in humanrobot interaction, where accurately interpreting human emotions through multiple modalities is essential for heartfelt communication. Although previous multimodal emotion recognition models have shown reasonable performance, there are two difficulties: (1) they often struggle to capture the fine-grained interactions between the modalities. (2) The prominent features are different across different emotions. To address these difficulties, we propose a novel Cross-modal Emotion-specific Attention model (CEA) for multimodal emotion recognition. The proposed model has two key components to address the difficulties above: (1) To capture the fine-grained interactions, we introduce the dense interaction matrix representation. (2) To focus more on emotion-specific features, we also introduce the specific emotion tokens. Combining these two components enhances the model’s ability to capture subtle emotional nuances and improves overall recognition accuracy. We evaluate the performance of the proposed architecture on the CREMA-D public audiovisual datasets through comprehensive ablation studies and comparison with baseline models. The results demonstrate that our Cross-modal Emotion-specific Attention model significantly outperforms the baseline methods, confirming its effectiveness in enhancing emotion recognition accuracy. Jia-Yi Chen, Vijay John, Yasutomo Kawanishi |
FG | 3 |
| 2025 | MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor FusionabstractMulti-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as widearea distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Vijay John, Takahiro Komamizu, Ichiro Ide |
FG | 2 |
| 2025 | FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D ImagesabstractThe ability to abstract complex 3D environments into simplified and structured representations is crucial across various domains. 3D semantic scene graphs (SSGs) achieve this by representing objects as nodes and their interrelationships as edges, facilitating high-level scene understanding. Existing methods for 3D SSG generation, however, face significant challenges, including high computational demands and non-incremental processing that hinder their suitability for real-time open-world applications. To address this issue, we propose FROSS (Faster-than-Real-Time Online 3D Semantic Scene Graph Generation), an innovative approach for online and faster-than-real-time 3D SSG generation that leverages the direct lifting of 2D scene graphs to 3D space and represents objects as 3D Gaussian distributions. This framework eliminates the dependency on precise and computationally-intensive point cloud processing. Furthermore, we extend the Replica dataset with inter-object relationship annotations, creating the ReplicaSSG dataset for comprehensive evaluation of FROSS. The experimental results from evaluations on ReplicaSSG and 3DSSG datasets show that FROSS can achieve superior performance while operating significantly faster than prior 3D SSG generation methods. Our implementation and dataset are publicly available at https://github.com/Howardkhh/FROSS. Hao-Yu Hou, Chun-Yi Lee, Motoharu Sonogashira, Yasutomo Kawanishi |
ICCV | 4 |
| 2025 | IntentVC 2025: The ACM Multimedia Grand Challenge on Intention-Oriented Controllable Video CaptioningabstractThe IntentVC Challenge, held in conjunction with ACM Multimedia 2025, introduces a novel benchmark for intention-oriented controllable video captioning. Unlike conventional captioning methods that generate generic, scene-level summaries, IntentVC focuses on intention-specific generation. Participants are required to produce captions explicitly conditioned on user-defined intentions, such as emphasizing a specific object tracked within a video. To support this task, the challenge provides an extended version of the LaSOT dataset annotated with intention-focused captions across 70 object categories. A standardized evaluation protocol and public leaderboard enable fair and reproducible comparison among submitted methods. By advancing research in personalized and adaptive video understanding, IntentVC offers a platform for exploring controllable vision-language modeling with practical relevance for accessibility, retrieval, and human-AI interaction. As a result, a total of 23 teams and 58 active participants have participated, and a total of 1,443 entries have been submitted. More information and resources are available at https://sites.google.com/view/intentvc/. Takahiro Komamizu, Marc A. Kastner 0001, Yasutomo Kawanishi, Trung Thanh Nguyen 0006, Junan Chen 0004 |
ACM Multimedia | 3 |
| 2025 | RoboDJ: Live Commentary Robots System Driven by Physical- and Cyber-World Observations
Yasutomo Kawanishi, Yutaka Nakamura, Taiken Shintani, Carlos Toshinori Ishi, Seiya Kawano, Koichiro Yoshino, Takashi Minato, Michihiko Minoh |
MMM (5) | 1 |
| 2025 | Towards Visual Storytelling by Understanding Narrative Context Through Scene-Graphs
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMM (4) | 3 |
| 2025 | Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing ModalitiesabstractDespite interest in multimodal classification, few studies have addressed the missing modality problem in which an incomplete multimodal input with one or more missing modalities is classified as the target class. The missing modality problem is shown to reduce the classification accuracy as the discriminative power of the obtained feature space is reduced. In this study, we address the missing modality problem in multimodal classification using a novel cascaded framework. The proposed framework is formulated in the feature space to address the missing modality problem by generating complete multimodal data from incomplete multimodal data. Subsequently, an optimal multimodal data is obtained by feature selection of the generated and original data. The proposed cascaded framework consists of three steps: feature extraction, feature generation, and classification. The framework is formulated to handle both complete and incomplete multimodal data simultaneously. The cascaded framework is trained using novel latent loss functions: missing modality joint loss, centroid joint loss, and latent prior loss. These loss functions, based on metric learning, are designed to ensure that data from the same class remain proximate in the latent space irrespective of the presence or absence of modality data. The cascaded framework is validated on bimodal audio-visible RAVDESS and trimodal audio-visible-thermal Speaking Faces datasets. The experimental results show that the cascaded framework improves classification accuracy even with incomplete multimodal data. Vijay John, Yasutomo Kawanishi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese QuestionsabstractSituated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in questions are often clarified by the contexts in conversational situations, such as joint attention with a user or user gaze information. In this study, we propose the Gaze-grounded VQA dataset (GazeVQA) that clarifies ambiguous questions using gaze information by focusing on a clarification process complemented by gaze information. We also propose a method that utilizes gaze target estimation results to improve the accuracy of GazeVQA tasks. Our experimental results showed that the proposed method improved the performance in some cases of a VQA system on GazeVQA and identified some typical problems of GazeVQA tasks that need to be improved. Shun Inadumi, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino |
LREC/COLING | 4 |
| 2024 | J-CRe3: A Japanese Conversation Dataset for Real-world Reference ResolutionabstractUnderstanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions to the visual information observed in egocentric views. To this end, we propose a multimodal reference resolution task and construct a Japanese Conversation dataset for Real-world Reference Resolution (J-CRe3). Our dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and an assistant robot at home. The dataset is annotated with crossmodal tags between phrases in the utterances and the object bounding boxes in the video frames. These tags include indirect reference relations, such as predicate-argument structures and bridging references as well as direct reference relations. We also constructed an experimental model and clarified the challenges in multimodal reference resolution tasks. Nobuhiro Ueda, Hideko Habe, Akishige Yuguchi, Seiya Kawano, Yasutomo Kawanishi, Sadao Kurohashi, Koichiro Yoshino |
LREC/COLING | 5 |
| 2024 | One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-Scale and Action Label FeaturesabstractOpen-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing and classifying actions based on a predefined set of categories. In contrast, Open-vocab TAD goes further and is not limited to these predefined categories. This is particularly useful in real-world scenarios where the variety of actions in videos can be vast and not always predictable. The prevalent methods in Open-vocab TAD typically employ a 2-stage approach, which involves generating action proposals and then identifying those actions. However, errors made during the first stage can adversely affect the subsequent action identification accuracy. Additionally, existing studies face challenges in handling actions of different durations owing to the use of fixed temporal processing methods. Therefore, we propose a L-stage approach consisting of two primary modules: Multi-scale Video Analysis (MVA) and Video-Text Alignment (VTA). The MVA module captures actions at varying temporal resolutions, overcoming the challenge of detecting actions with diverse durations. The VTA module leverages the synergy between visual and textual modalities to precisely align video segments with corresponding action labels, a critical step for accurate action identification in Open-vocab scenarios. Evaluations on widely recognized datasets THUMOSl4 and ActivityNet-I.3, showed that the proposed method achieved superior results compared to the other methods in both Open-vocab and Closed-vocab settings. This serves as a strong demonstration of the effectiveness of the proposed method in the TAD task. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
FG | 2 |
| 2024 | Recurrent Graph Convolutional Network for Sequential Pose Prediction from 3D Human Skeleton Sequence
Tomohiro Fujita, Yasutomo Kawanishi |
ICPR (14) | 2 |
| 2024 | Generating Pseudo-Strong Labels from Weak Labels for Distributed Multi-Microphone Sound Event Detection
Vijay John, Yasutomo Kawanishi |
ICPR (10) | 2 |
| 2024 | Action Selection Learning for Multi-label Multi-view Action RecognitionabstractMulti-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused on multi-view action recognition in a narrow area with strong labels available, where the onset and offset of each action are labeled at the frame-level. This study focuses on real-world scenarios where cameras are distributed to capture a wide-range area with only weak labels available at the video-level. We propose the method named Multi-view Action Selection Learning (MultiASL), which leverages action selection learning to enhance view fusion by selecting the most useful information from different viewpoints. The proposed method includes a Multi-view Spatial-Temporal Transformer video encoder to extract spatial and temporal features from multi-viewpoint videos. Action Selection Learning is employed at the frame-level, using pseudo ground-truth obtained from weak labels at the video-level, to identify the most relevant frames for action recognition. Experiments in a real-world office environment using the MM-Office dataset demonstrate the superior performance of the proposed method compared to existing methods. The source code is available at https://github.com/thanhhff/MultiASL/. Trung Thanh Nguyen 0006, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMAsia | 2 |
| 2024 | Computational measurement of perceived pointiness from pronunciationabstractAbstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications. Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 6 |
| 2024 | Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi |
Multim. Tools Appl. | 6 |
| 2023 | ManifoldNeRF: View-dependent Image Feature Supervision for Few-shot Neural Radiance Fields
Daiju Kanaoka, Motoharu Sonogashira, Hakaru Tamukoh, Yasutomo Kawanishi |
BMVC | 4 |
| 2023 | DeePoint: Visual Pointing Recognition and Direction EstimationabstractIn this paper, we realize automatic visual recognition and direction estimation of pointing. We introduce the first neural pointing understanding method based on two key contributions. The first is the introduction of a first-of-its-kind large-scale dataset for pointing recognition and direction estimation, which we refer to as the DP Dataset. DP Dataset consists of more than 2 million frames of 33 people pointing in various styles annotated for each frame with pointing timings and 3D directions. The second is DeePoint, a novel deep network model for joint recognition and 3D direction estimation of pointing. DeePoint is a Transformer-based network which fully leverages the spatio-temporal coordination of the body parts, not just the hands. Through extensive experiments, we demonstrate the accuracy and efficiency of DeePoint. We believe DP Dataset and DeePoint will serve as a sound foundation for visual human intention understanding. Shu Nakamura, Yasutomo Kawanishi, Shohei Nobuhara, Ko Nishino |
ICCV | 2 |
| 2023 | Operative Action Captioning for Estimating System ActionsabstractHuman-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express their understanding and action-planning results. In this study, we propose a new task of operative action captioning that estimates and verbalizes the actions to be taken by the system in a human-assisting domain. We constructed a system that outputs a verbal description of a possible operative action that changes the current state to the given target state. We collected a dataset consisting of two images as observations, which express the current state and the state changed by actions and a caption that describes the actions that change the current state to the target state, by crowdsourcing in daily life situations. Then we constructed a system that estimates an operative action by a caption. Since the operative action's caption is expected to contain some state-changing actions, we use scene graph prediction as an auxiliary task because the events written in the scene graphs correspond to the state changes. Experimental results showed that our system successfully described the operative actions that should be conducted between the current and target states. The auxiliary tasks that predict the scene graphs improved the quality of the estimation results. Taiki Nakamura, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino |
ICRA | 4 |
| 2023 | Audio-Visual Sensor Fusion Framework Using Person Attributes Robust to Missing Visual Modality for Person Recognition
Vijay John, Yasutomo Kawanishi |
MMM (2) | 2 |
| 2023 | Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide |
MMM (1) | 3 |
| 2023 | Multimodal Cascaded Framework with Metric Learning Robust to Missing Modalities for Person ClassificationabstractThis paper addresses the missing modality problem in multimodal person classification, where an incomplete multimodal input with one modality missing is classified into predefined person classes. A multimodal cascaded framework with three deep learning models is proposed, where model parameters, outputs, and latent space learnt at a given step are transferred to the model in a subsequent step. The cascaded framework addresses the missing modality problem by, firstly, generating the complete multimodal data from the incomplete multimodal data in the feature space via a latent space. Subsequently, the generated and original multimodal features are effectively merged and embedded into a final latent space to estimate the person label. During the learning phase, the cascaded framework uses two novel latent loss functions, the missing modality joint loss, and latent prior loss to learn the different latent spaces. The missing modality joint loss ensures that the similar class latent data are close to each other, even if a modality is missing. In the cascaded framework, the latent prior loss learns the final latent space using a previously learnt latent space as a prior. The proposed framework is validated on the audio-visible RAVDESS and the visible-thermal Speaking Faces datasets. A detailed comparative analysis and an ablation analysis are performed, which demonstrate that the proposed framework enhances the robustness of person classification even under conditions of missing modalities, reporting an average of 21.75% increase and 25.73% increase over the baseline algorithms on the RAVDESS and Speaking Faces datasets. Vijay John, Yasutomo Kawanishi |
MMSys | 2 |
| 2022 | Detection of Birds in a 3D Environment Referring to Audio-Visual InformationabstractWe propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed. Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi |
AVSS | 1 |
| 2022 | Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal StatesabstractThis paper presents an autonomous conversational mobile robot Butsukusa that can describe its own observations and internal states during patrolling tasks. The proposed robot can observe the surrounding environment using the recognition module for objects, humans, environment, localization, and speech and then move autonomously around an indoor living space. Interaction skills via language are required for the robot to perform in such human-centered spaces. To investigate a better communication protocol with users, we evaluate various language generation patterns based on different observations and interaction patterns. The evaluation results indicate that the importance of describing the robot's observation results and internal states, as well as the necessity of an appropriate description, depends on the situation. Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh |
HRI | 5 |
| 2022 | Audio and Video-based Emotion Recognition using Multimodal TransformersabstractEmotion recognition, an important research problem in human-robot interactions, is primarily achieved by extracting human emotions from audio and visual data. State-of-the-art performance is reported by audio-visual sensor fusion algorithms using deep learning models such as CNN, RNN, and LSTM. However, the RNN and LSTM are shown to be limited in handling the long-term dependencies over the entire input sequence. In this work, we propose to improve the performance of audio-visual emotion recognition using a novel transformer-based model, containing three transformer branches, named multimodal transformers. The three transformer branches, in our work, compute the audio self-attention, the video self-attention, and the audio-video cross attention. The self-attention branches identify the most relevant information in the audio and video input, while the cross-attention branch identifies the most relevant audio-video interactive information. The relevant information from these three branches report the best performance in our ablation study. We also propose a novel temporal embedding scheme, termed block embedding, to add the temporal information to the visual feature, derived from the multiple frames in the video. The proposed architecture is validated using the RAVDESS, CREMA-D, and SAVEE audio-visual public datasets. A detailed ablation study and comparative analysis with baseline models is performed. The results show that the proposed multi-modal transformer framework is better than the baseline methods. Vijay John, Yasutomo Kawanishi |
ICPR | 2 |
| 2022 | Label-based Multiple Object Ensemble Tracking with Randomized Frame DroppingabstractThis paper proposes a novel approach for ensemble tracking of multiple objects. Existing multiple object tracking methods often fail by appearance changes and occlusions of the targets, which lead to misdetection, mismatching, and drift of detections. The proposed method runs multiple "weak" trackers and aggregates the weak tracking results for each frame. Then, the proposed label-based ensemble is performed to track objects by considering a set of "weak" tracking results (instance IDs) for each target in a frame as a feature vector. This paper also proposes the randomized frame dropping that randomly drops input frames for each weak tracker to differentiate the inputs to the trackers. The proposed method can use any tracking algorithm as a weak tracker and pull up its performance. The performance of the proposed method has been confirmed by applying it to the current state-of-the-art method on MOT20, a public benchmark dataset. Yasutomo Kawanishi |
ICPR | 1 |
| 2022 | A Multimodal Sensor Fusion Framework Robust to Missing Modalities for Person RecognitionabstractUtilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal data is always available. In this paper, we propose a novel trimodal sensor fusion framework using the audio, visible, and thermal camera, which addresses the missing modality problem. In the framework, a novel deep latent embedding framework, termed the AVTNet, is proposed to learn multiple latent embeddings. Also, a novel loss function, termed missing modality loss, accounts for possible missing modalities based on the triplet loss calculation while learning the individual latent embeddings. Additionally, a joint latent embedding utilizing the trimodal data is learnt using the multi-head attention transformer, which assigns attention weights to the different modalities. The different latent embeddings are subsequently used to train a deep neural network. The proposed framework is validated on the Speaking Faces dataset. A comparative analysis with baseline algorithms shows that the proposed framework significantly increases the person recognition accuracy while accounting for missing modalities. Vijay John, Yasutomo Kawanishi |
MMAsia | 2 |
| 2021 | Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 4 |
| 2021 | Soft-Boundary Label Relaxation with class placement constraints for semantic segmentation of the railway environment
Yuki Furitsu, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Hiroki Mukojima, Nozomi Nagamine |
Pattern Recognit. Lett. | 3 |
| 2020 | LFIR2Pose: Pose Estimation from an Extremely Low-resolution FIR image SequenceabstractIn this paper, we propose a method for human pose estimation from a Low-resolution Far-InfraRed (LFIR) image sequence captured by a 16 × 16 FIR sensor array. Human body estimation from such a single LFIR image is a hard task. For training the estimation model, annotation of the human pose to the images is also a difficult task for human. Thus, we propose the LFIR2Pose model which accepts a sequence of LFIR images and outputs the human pose of the last frame, and also propose an automatic annotation system for the model training. Additionally, considering that the scale of human body motion is largely different among body parts, we also propose a loss function focusing on the difference. Through an experiment, we evaluated the human pose estimation accuracy with an original data set, and confirmed that human pose can be estimated accurately from an LFIR image sequence. Saki Iwata, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa |
ICPR | 2 |
| 2020 | Ω-GAN: Object Manifold Embedding GAN for Image Generation by Disentangling Parameters into Pose and Shape ManifoldsabstractIn this paper, we propose Object Manifold Embedding GAN (Ω-GAN) to generate images of variously shaped and arbitrarily posed objects from a noise variable sampled from a distribution defined over the pose and the shape manifolds in a vector space. We introduce Parametric Manifold Sampling to sample noise variables from a distribution over the pose manifold to conditionally generate object images in arbitrary poses by tuning the pose parameter. We also introduce Object Identity Loss for clearly disentangling the pose and shape parameters, which allows us to maintain the shape of the object instance when only the pose parameter is changed. Through evaluation, we confirmed that the proposed Ω-GAN could generate variously shaped object images in arbitrary poses by changing the pose and shape parameters independently. We also introduce an application of the proposed method for object pose estimation, through which we confirmed that the object poses in the generated images are accurate. Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
ICPR | 1 |
| 2020 | Median-Shape Representation Learning for Category-Level Object Pose Estimation in Cluttered EnvironmentsabstractIn this paper, we propose an occlusion-robust pose estimation method of an unknown object instance in an object category from a depth image. In a cluttered environment, objects are often occluded mutually. For estimating the pose of an object in such a situation, a method that de-occludes the unobservable area of the object would be effective. However, there are two difficulties; occlusion causes the offset between the center of the actual object and its observable area, and different instances in a category may have different shapes. To cope with these difficulties, we propose a two-stage Encoder-Decoder model to extract features with objects whose centers are aligned to the image center. In the model, we also propose the Median-shape Reconstructor as the second stage to absorb shape variations in a category. By evaluating the method with both a large-scale virtual dataset and a real dataset, we confirmed the proposed method achieves good performance on pose estimation of an occluded object from a depth image. Hiroki Tatemichi, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Ayako Amma, Hiroshi Murase |
ICPR | 2 |
| 2020 | Imageability Estimation using Visual and Language FeaturesabstractImageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability. Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase |
ICMR | 4 |
| 2020 | Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 3 |
| 2020 | More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase |
MMM (2) | 5 |
| 2020 | Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 4 |
| 2019 | Exemplar-Based Pseudo-Viewpoint Rotation for White-Cane User Recognition from a 2D Human Pose SequenceabstractIn recent years, various facilities are equipped to support visually impaired people, but accidents caused by visual disabilities still occur. In this paper, to support the visually-impaired people in a public space, we aim to classify whether a pedestrian image sequence obtained by a surveillance camera is a white-cane user or not from the temporal transition of a human pose represented as 2D coordinates. However, since the appearance of the 2D pose varies largely depending on the viewpoint of the pose, it is difficult to classify them. So, in this paper, we propose a method to rotate the viewpoint of a pose from various pseudo-viewpoints based on a pair of 2D poses simultaneously observed and classify the sequence by multiple classifiers corresponding to each viewpoint. Viewpoint rotation makes it possible to obtain pseudo-poses seen from various pseudo-viewpoints, extract richer pose features, and recognize white-cane users more accurately. Through an experiment, we confirmed that the proposed method improves the recognition rate by 12% compared to the method not employing viewpoint rotation. Naoki Nishida 0003, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Jun Piao |
AVSS | 2 |
| 2019 | Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase |
Multim. Tools Appl. | 3 |
| 2018 | Which Content in a Booklet is he/she Reading? Reading Content Estimation using an Indoor Surveillance CameraabstractIn this paper, we propose a method for estimating reading content in a booklet using an image captured by an indoor surveillance camera. Here, we assume that a reading content can be specified by estimating followings; what booklet, which page of the booklet, and which region in the page. We propose a reading booklet/page estimation method based on image search, and a reading region estimation method focusing on the body pose of the reader. We evaluated the method as a 44 classes classification problem, which consists of eleven pages of booklets and four regions in each pages. We achieved 25.6% in accuracy of the reading content estimation. Yasutomo Kawanishi, Hiroshi Murase, Kazuyuki Tasaka, Hiromasa Yanagihara |
ICPR | 1 |
| 2018 | Gaze-Inspired Learning for Estimating the Attractiveness of a Food PhotoabstractThe number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them. Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ISM | 4 |
| 2018 | Voting-based Hand-Waving Gesture Spotting from a Low-Resolution Far-Infrared Image SequenceabstractWe propose a temporal spotting method of a hand gesture from a low-resolution far-infrared image sequence captured by a far-infrared sensor array. The sensor array captures the spatial distribution of far-infrared intensity as a thermal image by detecting far-infrared waves emitted from heat sources. It is difficult to spot a hand gesture from a sequence of thermal images captured by the sensor due to its low-resolution, heavy noise, and varying duration of the gesture. Therefore, we introduce a voting-based approach to spot the gesture with template matching-based gesture recognition. We confirm the effectiveness of the proposed temporal spotting method in several settings. Yasutomo Kawanishi, Chisato Toriyama, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa, Masato Kawade |
VCIP | 1 |
| 2017 | Action recognition from extremely low-resolution thermal image sequenceabstractThis paper proposes a Deep Learning-based action recognition method from an extremely low-resolution thermal image sequence. The method recognizes daily actions by humans (e.g. walking, sitting down, standing up, etc.) and abnormal actions (e.g. falling down) without privacy concerns. While privacy concerns can be ignored, it is difficult to compute feature points and to obtain a clear edge of the human body from an extremely low-resolution thermal image. To address these problems, this paper proposes a Deep Learning-based action recognition method that combines convolution layers and an LSTM layer for learning spatio-temporal representation, whose inputs are the thermal images and their frame differences cropped by the gravity center of human regions. The effectiveness of the proposed method was confirmed through experiments. Takayuki Kawashima, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Daisuke Deguchi, Tomoyoshi Aizawa, Masato Kawade |
AVSS | 2 |
| 2017 | Automatic Selection of Web Contents Towards Automatic Authoring of a Video BiographyabstractIn this paper, we propose a method for image selection using Web image search for automatic video biography authoring. In the proposed method, images are selected from the image search results considering their visual contents for inclusion in the video biography. Through evaluation, we confirmed the effectiveness of the proposed image selection method compared to a baseline method which simply selects the top 1 search result. Ichiro Ide, Yasutomo Kawanishi, Kyoka Kunishiro, Frank Nack, Daisuke Deguchi, Hiroshi Murase |
ISM | 2 |
| 2017 | Summarization of News Videos Considering the Consistency of Auditory and Visual ContentsabstractSince news videos are valuable sources of multimedia information on real-world events, there is a demand for viewing them efficiently. However, there is a problem that summarization methods based on auditory contents do not take into account the visual contents. In the case of news videos, due to its presentation style where audio contents and visual contents do not necessarily come from the same source, this could severely decrease the amount of informative visual contents included in the generated summarized video. Thus, we propose a method for summarizing a sequence of news videos considering the consistency of both auditory and visual contents. The proposed method first selects key-sentences from the auditory contents (Closed Caption) of each news story in the sequence, and then selects a shot within the news story whose "Visual Concepts" detected from the visual contents are the most consistent with the key-phrase. Finally, the audio segment corresponding to each key-phrase is overlapped onto the selected shot, and then concatenated to generate a summarized video. The effectiveness of the proposed method was confirmed on several news topics through a subjective experiment. Ichiro Ide, Ryunosuke Tanishige, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase |
ISM | 5 |
| 2017 | Monocular localization within sparse voxel mapsabstractWe introduce a method that uses a single camera to localize a vehicle within a pre-constructed map consisting of a voxel occupancy grid and road-line marker positions. Sophisticated mapping hardware is capable of creating high-accuracy 3D maps of road environments, but localizing a vehicle within such maps is one of the challenges at the forefront of automated driving. A solution which is robust to dynamic environments, while using only inexpensive sensors, is a difficult problem. In addition, maps that enable precise localization consume a lot of data which is impractical for the expansive environments encountered in real-world road networks. We show how using the area of edge regions shared between rendered views of a compact voxel map and in-vehicle camera images can be coupled with non-linear optimization methods to determine the camera position and pose. David Wong 0002, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 2 |
| 2017 | Regression of feature scale tracklets for decimeter visual localization
David Wong 0002, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase |
Image Vis. Comput. | 3 |
| 2016 | A classification method of cooking operations based on eye movement patternsabstractWe are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables. Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase |
ETRA | 4 |
| 2016 | Moving camera background-subtraction for obstacle detection on railway tracksabstractWe propose a method for detecting obstacles by comparing input and reference train frontal view camera images. In the field of obstacle detection, most methods employ a machine learning approach, so they can only detect pre-trained classes, such as pedestrian, bicycle, etc. This means that obstacles of unknown classes cannot be detected. To overcome this problem, we propose a background subtraction method that can be applied to moving cameras. First, the proposed method computes frame-by-frame correspondences between the current and the reference (database) image sequences. Then, obstacles are detected by applying image subtraction to corresponding frames. To confirm the effectiveness of the proposed method, we conducted an experiment using several image sequences captured on an experimental track. Its results showed that the proposed method could detect various obstacles accurately and effectively. Hiroki Mukojima, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Masato Ukai, Nozomi Nagamine, Ryuta Nakasone |
ICIP | 3 |
| 2016 | Misclassification tolerable learning for robust pedestrian orientation classificationabstractIn this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset. Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
ICPR | 1 |
| 2016 | Parts Selective DPM for detection of pedestrians possessing an umbrellaabstractIn recent years, pedestrian detection from an in-vehicle camera has been attracting attention. However, in the case of a raining situation, the detection accuracy decreases because the head of a pedestrian tends to be occluded by an umbrella. In oder to handle such cases, in this paper, as a variation of the Deformable Part Model (DPM) which is widely used in the field of object recognition, we propose “Parts Selective DPM (PS-DPM)” which selectively chooses the original part filters and additional part filters trained independently. In the detection of pedestrians possessing an umbrella, the selection of head and umbrella parts will make pedestrian detection more robust to the occlusion. We conducted experiments to evaluate the performance of the proposed method. As a result, pedestrian detection with the proposed PS-DPM achieved high detection accuracy in rainy weather, compared with the detection by the conventional DPM. Moreover, we confirmed that it did not decrease the pedestrian detection accuracy in fine weather. Yuto Shimbo, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase |
Intelligent Vehicles Symposium | 2 |
| 2015 | Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF cameraabstractThis paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification. Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi |
Intelligent Vehicles Symposium | 2 |
| 2010 | Privacy-Protected Camera for the Sensing Web
Ikuhisa Mitsugami, Masayuki Mukunoki, Yasutomo Kawanishi, Hironori Hattori, Michihiko Minoh |
IPMU (2) | 3 |
| 2009 | Background Estimation Based on Device Pixel Structures for Silhouette Extraction
Yasutomo Kawanishi, Takuya Funatomi, Koh Kakusho, Michihiko Minoh |
ACCV (3) | 1 |