Tatsuya Harada

dblp:14/5849 · DBLP profile ↗
← Back
189ranked-venue papers
13as first author
78since 2021 · last 2026
0000-0002-3712-3691ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 139 · 12 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 110 · 3 first-author · 45 since 2021Systems, architecture and hardware · 36 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
abstract
This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rigorous filtering/deduplication, object-detection-driven evidence extraction, and Large Language Model (LLM)-based refinement under grounding constraints. Using this pipeline, we build two resources: an image-caption dataset (DEJIMA-Cap) and a VQA dataset (DEJIMA-VQA), each containing 3.88M image-text pairs, far exceeding the size of existing Japanese V&L datasets. Human evaluations demonstrate that DEJIMA achieves substantially higher Japaneseness and linguistic naturalness than datasets constructed via translation or manual annotation, while maintaining factual correctness at a level comparable to human-annotated corpora. Quantitative analyses of image feature distributions further confirm that DEJIMA broadly covers diverse visual domains characteristic of Japan, complementing its linguistic and cultural representativeness. Models trained on DEJIMA exhibit consistent improvements across multiple Japanese multimodal benchmarks, confirming that culturally grounded, large-scale resources play a key role in enhancing model performance. All data sources and modules in our pipeline are licensed for commercial use, and we publicly release the resulting dataset and metadata to encourage further research and industrial applications in Japanese V&L modeling.
Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando, Yusuke Mukuta, Kohei Uehara, Tatsuya Harada
LREC6
2026 SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
abstract
Grounding complex, compositional visual queries with multiple objects and relationships is a fundamental challenge for vision-language models. While standard phrase grounding methods excel at localizing single objects, they lack the structural inductive bias to parse intricate relational descriptions, often failing as queries become more descriptive. To address this structural deficit, we focus on scenegraph grounding, a powerful but less-explored formulation where the query is an explicit graph of objects and their relationships. However, existing methods for this task also struggle, paradoxically showing decreased performance as the query graph grows—failing to leverage the very information that should make grounding easier. We introduce SceneProp, a novel method that resolves this issue by reformulating scene-graph grounding as a Maximum a Posteriori (MAP) inference problem in a Markov Random Field (MRF). By performing global inference over the entire query graph, SceneProp finds the optimal assignment of image regions to nodes that jointly satisfies all constraints. This is achieved within an end-to-end framework via a differentiable implementation of the Belief Propagation algorithm. Experiments on four benchmarks show that our dedicated focus on the scene-graph grounding formulation allows SceneProp to significantly outperform prior work. Critically, its accuracy consistently improves with the size and complexity of the query graph, demonstrating for the first time that more relational context can, and should, lead to better grounding. Codes are available at https://github.com/keitaotani/SceneProp.
Keita Otani, Tatsuya Harada
WACV2
2026 Semi-supervised Domain Adaptation via Mutual Alignment through Joint Error
abstract
Most existing methods for unsupervised domain adaptation focus on learning domain-invariant representations. However, recent works have shown that the generalization on the target domain can fail due to the trade-off between marginal distribution alignment and joint error under a large domain shift. A few labeled target data points can enhance adaptation quality, but the distribution shift between labeled and unlabeled target data is often overlooked. Therefore, we propose a novel learning theory to address the joint error in semi-supervised domain adaptation that can reduce the mutual distribution shift between pairs from labeled and unlabeled domains. Furthermore, we introduce a discrepancy measurement between hypotheses to tackle the inconsistency of the loss functions in the algorithm and theory. Extensive experiments demonstrate that our method consistently outperforms baseline approaches, particularly in scenarios with large domain shifts and scarce labeled target data.
Dexuan Zhang, Thomas Westfechtel, Tatsuya Harada
WACV3
2025 Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustment
abstract
Capturing high-quality photographs under diverse real-world lighting conditions is challenging, as both natural lighting (e.g., low-light) and camera exposure settings (e.g., exposure time) significantly impact image quality. This challenge becomes more pronounced in multi-view scenarios, where variations in lighting and image signal processor (ISP) settings across viewpoints introduce photometric inconsistencies. Such lighting degradations and view-dependent variations pose substantial challenges to novel view synthesis (NVS) frameworks based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS).To address this, we introduce Luminance-GS, a novel approach to achieving high-quality novel view synthesis results under diverse challenging lighting conditions using 3DGS. By adopting per-view color matrix mapping and view adaptive curve adjustments, Luminance-GS achieves state-of-the-art (SOTA) results across various lighting conditions—including low-light, overexposure, and varying exposure—while not altering the original 3DGS explicit representation. Compared to previous NeRF- and 3DGS-based baselines, Luminance-GS provides real-time rendering speed with improved reconstruction quality. The source code is available at1.
Ziteng Cui, Xuangeng Chu, Tatsuya Harada
CVPR3
2025 A Theory of Learning Unified Model via Knowledge Integration from Label Space Varying Domains
abstract
Existing domain adaptation systems can hardly be applied to real-world problems with new classes presenting at deployment time, especially regarding source-free scenarios where multiple source domains do not share the label space despite being given a few labeled target data. To address this, we consider a challenging problem: multi-source semi-supervised open-set domain adaptation and propose a learning theory via joint error, effectively tackling strong domain shift. To generalize the algorithm into source-free cases, we introdcue a computationally efficient and architecture-flexible attention-based feature generation module. Extensive experiments on various data sets demonstrate the significant improvement of our proposed algorithm over baselines.
Dexuan Zhang, Thomas Westfechtel, Tatsuya Harada
CVPR3
2025 T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask Learning
abstract
We introduce T2V2 (**T**ext to **V**oice and **V**oice to **T**ext), a unified non-autoregressive model capable of performing both automatic speech recognition (ASR) and text-to-speech (TTS) synthesis within the same framework. T2V2 uses a shared Conformer backbone with rotary positional embeddings to efficiently handle these core tasks, with ASR trained using Connectionist Temporal Classification (CTC) loss and TTS using masked language modeling (MLM) loss. The model operates on discrete tokens, where speech tokens are generated by clustering features from a self-supervised learning model. To further enhance performance, we introduce auxiliary tasks: CTC error correction to refine raw ASR outputs using contextual information from speech embeddings, and unconditional speech MLM, enabling classifier free guidance to improve TTS. Our method is self-contained, leveraging intermediate CTC outputs to align text and speech using Monotonic Alignment Search, without relying on external aligners. We perform extensive experimental evaluation to verify the efficacy of the T2V2 framework, achieving state-of-the-art performance on TTS task and competitive performance in discrete ASR.
Nabarun Goswami, Hanqin Wang, Tatsuya Harada
ICLR3
2025 Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
abstract
For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q-values for the current policy using the Bellman operator. These algorithms for continuous actions rely exclusively on policy updates for improvement, which often results in low sample efficiency. This study examines the effectiveness of incorporating the Bellman optimality operator into actor-critic frameworks. Experiments in a simple environment show that modeling optimal values accelerates learning but leads to overestimation bias. To address this, we propose an annealing approach that gradually transitions from the Bellman optimality operator to the Bellman operator, thereby accelerating learning while mitigating bias. Our method, combined with TD3 and SAC, significantly outperforms existing approaches across various locomotion and manipulation tasks, demonstrating improved performance and robustness to hyperparameters related to optimality. The code for this study is available at https://github.com/motokiomura/annealed-q-learning.
Motoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada
ICML5
2025 FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge
Nabarun Goswami, Tatsuya Harada
INTERSPEECH2
2025 Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task Conditioning
abstract
We introduce Dr. RAW, a unified and tuning-efficient framework for high-level computer vision tasks directly operating on camera RAW data. Unlike previous approaches that optimize image signal processing (ISP) pipelines and fully fine-tune networks for each task, Dr. RAW achieves state-of-the-art performance with minimal parameter updates. At the input stage, we apply lightweight pre-processing modules, sensor and illumination mapping, followed by re-mosaicing, to mitigate data inconsistencies stemming from sensor variation and lighting. At the network level, we introduce task-specific adaptation through two modules: Sensor Prior Prompts (SPP) and Low-Rank Adaptation (LoRA). SPP injects sensor-aware conditioning into the network via learnable prompts derived from imaging priors, while LoRA enables efficient task-specific tuning by updating only low-rank matrices in key backbone layers. Despite minimal tuning, our method delivers superior results across four RAW-based tasks (object detection, semantic segmentation, instance segmentation, and pose estimation) on nine datasets encompassing low-light and over-exposed conditions. By harnessing the intrinsic physical cues of RAW data alongside parameter-efficient techniques, our method advances RAW-based vision systems, achieving both high accuracy and computational economy. We will release our source code.
Wenjun Huang 0001, Ziteng Cui, Yinqiang Zheng, Yirui He, Tatsuya Harada, Mohsen Imani
NeurIPS5
2025 I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
abstract
Participating in efforts to endow generative AI with the 3D physical world perception, we propose I2-NeRF, a novel neural radiance field framework that enhances isometric and isotropic metric perception under media degradation. While existing NeRF models predominantly rely on object-centric sampling, I2-NeRF introduces a reverse-stratified upsampling strategy to achieve near-uniform sampling across 3D space, thereby preserving isometry. We further present a general radiative formulation for media degradation that unifies emission, absorption, and scattering into a particle model governed by the Beer–Lambert attenuation law. By matting direct and media-induced in-scatter radiance, this formulation extends naturally to complex media environments such as underwater, haze, and even low-light scenes. By treating light propagation uniformly in both vertical and horizontal directions, I2-NeRF enables isotropic metric perception and can even estimate medium properties such as water depth. Experiments on real-world datasets demonstrate that our method significantly improves both reconstruction fidelity and physical plausibility compared to existing approaches. The source code is available at https://github.com/ShuhongLL/I2-NeRF.
Shuhong Liu, Lin Gu 0003, Ziteng Cui, Xuangeng Chu, Tatsuya Harada
NeurIPS5
2025 Intend to Move: A Multimodal Dataset for Intention-Aware Human Motion Understanding
abstract
Human motion is inherently intentional, yet most motion modeling paradigms focus on low-level kinematics, overlooking the semantic and causal factors that drive behavior. Existing datasets further limit progress: they capture short, decontextualized actions in static scenes, providing little grounding for embodied reasoning. To address these limitations, we introduce $\textit{Intend to Move (I2M)}$, a large-scale, multimodal dataset for intention-grounded motion modeling. I2M contains 10.1 hours of two-person 3D motion sequences recorded in dynamic realistic home environments, accompanied by multi-view RGB-D video, 3D scene geometry, and language annotations of each participant’s evolving intentions. Benchmark experiments reveal a fundamental gap in current motion models: they fail to translate high-level goals into physically and socially coherent motion. I2M thus serves not only as a dataset but as a benchmark for embodied intelligence, enabling research on models that can reason about, predict, and act upon the ``why'' behind human motion.
Ryo Umagami, Liu Yue, Xuangeng Chu, Ryuto Fukushima, Tetsuya Narita, Yusuke Mukuta, Tomoyuki Takahata, Tatsuya Harada
NeurIPS9
2025 ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model
abstract
Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow generation speed limits their application potential. In this paper, we introduce a novel autoregressive model that achieves real-time generation of highly synchronized lip movements and realistic head poses and eye blinks by learning a mapping from speech to a multi-scale motion codebook. Furthermore, our model can adapt to unseen speaking styles, enabling the creation of 3D talking avatars with unique personal styles beyond the identities seen during training. Extensive evaluations and user studies demonstrate that our method outperforms existing approaches in lip synchronization accuracy and perceived quality. Demos and codes are available at https://xg-chu.site/project_artalk/.
Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, Tatsuya Harada
SIGGRAPH Asia5
2025 Physiology-Aware PolySnake for Coronary Vessel Segmentation
abstract
Coronary artery disease (CAD) is a significant health risk that requires early detection for effective treatment. While recent advances in deep learning have shown promise in automating CAD detection from coronary computed to-mography angiography (CCTA) images, the accurate segmentation of coronary vessels remains a challenge, particularly due to the imbalanced presence of plaque in unhealthy vessels. This paper introduces a physiology-aware approach11https://github.com/opensourcetorch/Physiology-aware-PolySnake to coronary vessel segmentation that addresses these challenges. Our proposed pipeline consists of three main components. First, a hybrid UNeXt architecture is designed to segment artery boundaries and predict initial boundary contours by leveraging 3D spatial relations among adjacent slices. Second, we introduce multi-class circular convolution for iterative contour deformation, which generates well-connected contour pairs of the artery wall's inner and outer boundaries through iterative refinement. Finally, we propose a focal smooth Lllossfunction to handle the implicit class imbalance caused by plaque in unhealthy vessels and to enhance the robustness of the physiology-aware polysnake network by explicitly limiting the accuracy of initial contours. Extensive evaluations demonstrate that our methods significantly improve model performance, achieving state-of-the-art results in coronary vessel segmentation.
Yizhe Ruan, Lin Gu 0003, Yusuke Kurose, Junichi Iho, Youji Tokunaga, Makoto Horie, Yusaku Hayashi, Keisuke Nishizawa, Yasushi Koyama, Tatsuya Harada
WACV10
2025 Combining Inherent Knowledge of Vision-Language Models with Unsupervised Domain Adaptation Through Strong-Weak Guidance
abstract
Unsupervised domain adaptation (UDA) tries to overcome the tedious work of labeling data by leveraging a labeled source dataset and transferring its knowledge to a similar but different target dataset. Meanwhile, current vision-language models exhibit remarkable zero-shot prediction capabilities. In this work, we combine knowledge gained through UDA with the inherent knowledge of vision-language models. We introduce a strong-weak guidance learning scheme that employs zero-shot predictions to help align the source and target dataset. For the strong guidance, we expand the source dataset with the most confident samples of the target dataset. Additionally, we employ a knowledge distillation loss as weak guidance. The strong guidance uses hard labels but is only applied to the most confident predictions from the target dataset. Conversely, the weak guidance is employed to the whole dataset but uses soft labels. The weak guidance is implemented as a knowledge distillation loss with (adjusted) zero-shot predictions. We show that our method complements and benefits from prompt adaptation techniques for vision-language models. We conduct experiments and ablation studies on three benchmarks (OfficeHome, VisDA, and DomainNet), outperforming state-of-the-art methods. Our ablation studies further demonstrate the contributions of different components of our algorithm.
Thomas Westfechtel, Dexuan Zhang, Tatsuya Harada
WACV3
2025 Defender of privacy and fairness: Tiny but reversible generative model via mutually collaborative knowledge distillation
Sissi Xiaoxiao Wu, Zehong Huang, Zhicong Liang, Lin Gu 0003, Tatsuya Harada, Yingying Zhu 0004
Neurocomputing5
2025 A New Benchmark: Clinical Uncertainty and Severity Aware Labeled Chest X-Ray Images With Multi-Relationship Graph Learning
abstract
Chest radiography, commonly known as CXR, is frequently utilized in clinical settings to detect cardiopulmonary conditions. However, even seasoned radiologists might offer different evaluations regarding the seriousness and uncertainty associated with observed abnormalities. Previous research has attempted to utilize clinical notes to extract abnormal labels for training deep-learning models in CXR image diagnosis. However, these methods often neglected the varying degrees of severity and uncertainty linked to different labels. In our study, we initially assembled a comprehensive new dataset of CXR images based on clinical textual data, which incorporated radiologists' assessments of uncertainty and severity. Using this dataset, we introduced a multi-relationship graph learning framework that leverages spatial and semantic relationships while addressing expert uncertainty through a dedicated loss function. Our research showcases a notable enhancement in CXR image diagnosis and the interpretability of the diagnostic model, surpassing existing state-of-the-art methodologies. The dataset address of disease severity and uncertainty we extracted is: https://physionet.org/content/cad-chest/1.0/.
Mengliang Zhang, Xinyue Hu 0002, Lin Gu 0003, Kazuma Kobayashi, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0004
IEEE Trans. Medical Imaging6
2024 Aleth-NeRF: Illumination Adaptive NeRF with Concealing Field Assumption
abstract
The standard Neural Radiance Fields (NeRF) paradigm employs a viewer-centered methodology, entangling the aspects of illumination and material reflectance into emission solely from 3D points. This simplified rendering approach presents challenges in accurately modeling images captured under adverse lighting conditions, such as low light or over-exposure. Motivated by the ancient Greek emission theory that posits visual perception as a result of rays emanating from the eyes, we slightly refine the conventional NeRF framework to train NeRF under challenging light conditions and generate normal-light condition novel views unsupervisedly. We introduce the concept of a ``Concealing Field," which assigns transmittance values to the surrounding air to account for illumination effects. In dark scenarios, we assume that object emissions maintain a standard lighting level but are attenuated as they traverse the air during the rendering process. Concealing Field thus compel NeRF to learn reasonable density and colour estimations for objects even in dimly lit situations. Similarly, the Concealing Field can mitigate over-exposed emissions during rendering stage. Furthermore, we present a comprehensive multi-view dataset captured under challenging illumination conditions for evaluation. Our code and proposed dataset are available at https://github.com/cuiziteng/Aleth-NeRF.
Ziteng Cui, Lin Gu 0003, Xiao Sun 0001, Xianzheng Ma, Yu Qiao 0001, Tatsuya Harada
AAAI6
2024 Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
abstract
In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error distribution. However, a recent study suggested that the error distribution for training the value function is often skewed because of the properties of the Bellman operator, and violates the implicit assumption of normal error distribution in the least squares method. To address this, we proposed a method called Symmetric Q-learning, in which the synthetic noise generated from a zero-mean distribution is added to the target values to generate a Gaussian error distribution. We evaluated the proposed method on continuous control benchmark tasks in MuJoCo. It improved the sample efficiency of a state-of-the-art reinforcement learning method by reducing the skewness of the error distribution.
Motoki Omura, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada
AAAI4
2024 Discovering an Image-Adaptive Coordinate System for Photography Processing
Ziteng Cui, Lin Gu 0003, Tatsuya Harada
BMVC3
2024 RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images
Ziteng Cui, Tatsuya Harada
ECCV (4)2
2024 Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
Djamahl Etchegaray, Zi Huang, Tatsuya Harada, Yadan Luo
ECCV (40)3
2024 Open-Set Domain Adaptation via Joint Error Based Multi-class Positive and Unlabeled Learning
Dexuan Zhang, Thomas Westfechtel, Tatsuya Harada
ECCV (73)3
2024 GPAvatar: Generalizable and Precise Head Avatar from Image(s)
abstract
Head avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expressions and postures. Existing methods, categorized into 2D-based warping, mesh-based, and neural rendering approaches, present challenges in maintaining multi-view consistency, incorporating non-facial information, and generalizing to new identities. In this paper, we propose a framework named GPAvatar that reconstructs 3D head avatars from one or several images in a single forward pass. The key idea of this work is to introduce a dynamic point-based expression field driven by a point cloud to precisely and effectively capture expressions. Furthermore, we use a Multi Tri-planes Attention (MTA) fusion module in tri-planes canonical field to leverage information from multiple input images. The proposed method achieves faithful identity reconstruction, precise expression control, and multi-view consistency, demonstrating promising results for free-viewpoint rendering and novel view synthesis.
Xuangeng Chu, Ailing Zeng, Lijian Lin, Tatsuya Harada
ICLR7
2024 Discovering Multiple Solutions from a Single Task in Offline Reinforcement Learning
abstract
Recent studies on online reinforcement learning (RL) have demonstrated the advantages of learning multiple behaviors from a single task, as in the case of few-shot adaptation to a new environment. Although this approach is expected to yield similar benefits in offline RL, appropriate methods for learning multiple solutions have not been fully investigated in previous studies. In this study, we therefore addressed the problem of finding multiple solutions from a single task in offline RL. We propose algorithms that can learn multiple solutions in offline RL, and empirically investigate their performance. Our experimental results show that the proposed algorithm learns multiple qualitatively and quantitatively distinctive solutions in offline RL.
Takayuki Osa, Tatsuya Harada
ICML2
2024 Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration
abstract
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io.
Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin
ICRA229
2024 Robustifying a Policy in Multi-Agent RL with Diverse Cooperative Behaviors and Adversarial Style Sampling for Assistive Tasks
abstract
Autonomous assistance of people with motor impairments is one of the most promising applications of autonomous robotic systems. Recent studies have reported encouraging results using deep reinforcement learning (RL) in the healthcare domain. Previous studies showed that assistive tasks can be formulated as multi-agent RL, wherein there are two agents: a caregiver and a care-receiver. However, policies trained in multi-agent RL are often sensitive to the policies of other agents. In such a case, a trained caregiver’s policy may not work for different care-receivers. To alleviate this issue, we propose a framework that learns a robust caregiver’s policy by training it for diverse care-receiver responses. In our framework, diverse care-receiver responses are autonomously learned through trials and errors. In addition, to robustify the care-giver’s policy, we propose a strategy for sampling a care-receiver’s response in an adversarial manner during the training. We evaluated the proposed method using tasks in an Assistive Gym. We demonstrate that policies trained with a popular deep RL method are vulnerable to changes in policies of other agents and that the proposed framework improves the robustness against such changes.
Takayuki Osa, Tatsuya Harada
ICRA2
2024 Generalizable and Animatable Gaussian Head Avatar
abstract
In this paper, we propose Generalizable and Animatable Gaussian head Avatar (GAGA) for one-shot animatable head avatar reconstruction. Existing methods rely on neural radiance fields, leading to heavy rendering consumption and low reenactment speeds. To address these limitations, we generate the parameters of 3D Gaussians from a single image in a single forward pass. The key innovation of our work is the proposed dual-lifting method, which produces high-fidelity 3D Gaussians that capture identity and facial details. Additionally, we leverage global image features and the 3D morphable model to construct 3D Gaussians for controlling expressions. After training, our model can reconstruct unseen identities without specific optimizations and perform reenactment rendering at real-time speeds. Experiments show that our method exhibits superior performance compared to previous methods in terms of reconstruction quality and expression accuracy. We believe our method can establish new benchmarks for future research and advance applications of digital avatars.
Xuangeng Chu, Tatsuya Harada
NeurIPS2
2024 Style-NeRF2NeRF: 3D Style Transfer from Style-Aligned Multi-View Images
Haruo Fujiwara, Yusuke Mukuta, Tatsuya Harada
SIGGRAPH Asia3
2024 Soft Curriculum for Learning Conditional GANs with Noisy-Labeled and Uncurated Unlabeled Data
abstract
Label-noise or curated unlabeled data are used to compensate for the assumption of clean labeled data in training the conditional generative adversarial network; however, satisfying such an extended assumption is occasionally laborious or impractical. As a step towards generative modeling accessible to everyone, we introduce a novel conditional image generation framework that accepts noisy-labeled and uncurated unlabeled data during training: (i) closed-set and open-set label noise in labeled data and (ii) closed-set and open-set unlabeled data. To combat it, we propose soft curriculum learning, which assigns instance-wise weights for adversarial training while assigning new labels for unlabeled data and correcting wrong labels for labeled data. Unlike popular curriculum learning, which uses a threshold to pick the training samples, our soft curriculum controls the effect of each training instance by using the weights predicted by the auxiliary classifier, resulting in the preservation of useful samples while ignoring harmful ones. Our experiments show that our approach outperforms existing semi-supervised and label-noise robust methods in terms of both quantitative and qualitative performance. In particular, the proposed approach matches the performance of (semi-)supervised GANs even with less than half the labeled data.1
Kai Katsumata, Duc Minh Vo, Tatsuya Harada, Hideki Nakayama
WACV3
2024 Gradual Source Domain Expansion for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) tries to overcome the need for a large labeled dataset by transferring knowledge from a source dataset, with lots of labeled data, to a target dataset, that has no labeled data. Since there are no labels in the target domain, early misalignment might propagate into the later stages and lead to an error build-up. In order to overcome this problem, we propose a gradual source domain expansion (GSDE) algorithm. GSDE trains the UDA task several times from scratch, each time reinitializing the network weights, but each time expands the source dataset with target data. In particular, the highest-scoring target data of the previous run are employed as pseudo-source samples with their respective pseudo-label. Using this strategy, the pseudo-source samples induce knowledge extracted from the previous run directly from the start of the new training. This helps align the two domains better, especially in the early training epochs. In this study, we first introduce a strong baseline network and apply our GSDE strategy to it. We conduct experiments and ablation studies on three benchmarks (Office-31, OfficeHome, and DomainNet) and outperform state-of-the-art methods. We further show that the proposed GSDE strategy can improve the accuracy of a variety of different state-of-the-art UDA approaches.
Thomas Westfechtel, Hao-Wei Yeh, Dexuan Zhang, Tatsuya Harada
WACV4
2024 Can physician judgment enhance model trustworthiness? A case study on predicting pathological lymph nodes in rectal cancer
abstract
Explainability is key to enhancing the trustworthiness of artificial intelligence in medicine. However, there exists a significant gap between physicians' expectations for model explainability and the actual behavior of these models. This gap arises from the absence of a consensus on a physician-centered evaluation framework, which is needed to quantitatively assess the practical benefits that effective explainability should offer practitioners. Here, we hypothesize that superior attention maps, as a mechanism of model explanation, should align with the information that physicians focus on, potentially reducing prediction uncertainty and increasing model reliability. We employed a multimodal transformer to predict lymph node metastasis of rectal cancer using clinical data and magnetic resonance imaging. We explored how well attention maps, visualized through a state-of-the-art technique, can achieve agreement with physician understanding. Subsequently, we compared two distinct approaches for estimating uncertainty: a standalone estimation using only the variance of prediction probability, and a human-in-the-loop estimation that considers both the variance of prediction probability and the quantified agreement. Our findings revealed no significant advantage of the human-in-the-loop approach over the standalone one. In conclusion, this case study did not confirm the anticipated benefit of the explanation in enhancing model reliability. Superficial explanations could do more harm than good by misleading physicians into relying on uncertain predictions, suggesting that the current state of attention mechanisms should not be overestimated in the context of model explainability.
Kazuma Kobayashi, Yasuyuki Takamizawa, Mototaka Miyake, Sono Ito, Lin Gu 0003, Tatsuya Nakatsuka, Yu Akagi, Tatsuya Harada, Yukihide Kanemitsu, Ryuji Hamamoto
Artif. Intell. Medicine8
2024 Learning by Asking Questions for Knowledge-Based Novel Object Recognition
abstract
Abstract In real-world object recognition, there are numerous object classes to be recognized. Traditional image recognition methods based on supervised learning can only recognize object classes present in the training data, and have limited applicability in the real world. In contrast, humans can recognize novel objects by questioning and acquiring knowledge about them. Inspired by this, we propose a framework for acquiring external knowledge by generating questions that enable the model to instantly recognize novel objects. Our framework comprises three components: the object classifier (OC), which performs knowledge-based object recognition, the question generator (QG), which generates knowledge-aware questions to acquire novel knowledge, and the policy decision (PD) Model, which determines the “policy” of questions to be asked. The PD model utilizes two strategies, namely “confirmation” and “exploration”—the former confirms candidate knowledge while the latter explores completely new knowledge. Our experiments demonstrate that the proposed pipeline effectively acquires knowledge about novel objects compared to several baselines, and realizes novel object recognition utilizing the obtained knowledge. We also performed a real-world evaluation in which humans responded to the generated questions, and the model used the acquired knowledge to retrain the OC, which is a fundamental step toward a real-world human-in-the-loop learning-by-asking framework. We plan to release the dataset immediately upon acceptance of our work.
Kohei Uehara, Tatsuya Harada
Int. J. Comput. Vis.2
2024 Interpretable medical image Visual Question Answering via multi-modal relationship graph learning
Xinyue Hu 0002, Lin Gu 0003, Kazuma Kobayashi, Mengliang Zhang, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0003
Medical Image Anal.6
2024 Sketch-based semantic retrieval of medical images
abstract
The volume of medical images stored in hospitals is rapidly increasing; however, the utilization of these accumulated medical images remains limited. Existing content-based medical image retrieval (CBMIR) systems typically require example images, leading to practical limitations, such as the lack of customizable, fine-grained image retrieval, the inability to search without example images, and difficulty in retrieving rare cases. In this paper, we introduce a sketch-based medical image retrieval (SBMIR) system that enables users to find images of interest without the need for example images. The key concept is feature decomposition of medical images, which allows the entire feature of a medical image to be decomposed into and reconstructed from normal and abnormal features. Building on this concept, our SBMIR system provides an easy-to-use two-step graphical user interface: users first select a template image to specify a normal feature and then draw a semantic sketch of the disease on the template image to represent an abnormal feature. The system integrates both types of input to construct a query vector and retrieves reference images. For evaluation, ten healthcare professionals participated in a user test using two datasets. Consequently, our SBMIR system enabled users to overcome previous challenges, including image retrieval based on fine-grained image characteristics, image retrieval without example images, and image retrieval for rare cases. Our SBMIR system provides on-demand, customizable medical image retrieval, thereby expanding the utility of medical image databases.
Kazuma Kobayashi, Lin Gu 0003, Ryuichiro Hataya, Takaaki Mizuno, Mototaka Miyake, Hirokazu Watanabe, Masamichi Takahashi, Yasuyuki Takamizawa, Yukihiro Yoshida, Nobuji Kouno, Amina Bolatkan, Yusuke Kurose, Tatsuya Harada, Ryuji Hamamoto
Medical Image Anal.14
2024 Rethinking masked image modelling for medical image representation
abstract
Masked Image Modelling (MIM), a form of self-supervised learning, has garnered significant success in computer vision by improving image representations using unannotated data. Traditional MIMs typically employ a strategy of random sampling across the image. However, this random masking technique may not be ideally suited for medical imaging, which possesses distinct characteristics divergent from natural images. In medical imaging, particularly in pathology, disease-related features are often exceedingly sparse and localized, while the remaining regions appear normal and undifferentiated. Additionally, medical images frequently accompany reports, directly pinpointing pathological changes' location. Inspired by this, we propose Masked medical Image Modelling (MedIM), a novel approach, to our knowledge, the first research that employs radiological reports to guide the masking and restore the informative areas of images, encouraging the network to explore the stronger semantic representations from medical images. We introduce two mutual comprehensive masking strategies, knowledge-driven masking (KDM), and sentence-driven masking (SDM). KDM uses Medical Subject Headings (MeSH) words unique to radiology reports to identify symptom clues mapped to MeSH words (e.g., cardiac, edema, vascular, pulmonary) and guide the mask generation. Recognizing that radiological reports often comprise several sentences detailing varied findings, SDM integrates sentence-level information to identify key regions for masking. MedIM reconstructs images informed by this masking from the KDM and SDM modules, promoting a comprehensive and enriched medical image representation. Our extensive experiments on seven downstream tasks covering multi-label/class image classification, pneumothorax segmentation, and medical image-report analysis, demonstrate that MedIM with report-guided masking achieves competitive performance. Our method substantially outperforms ImageNet pre-training, MIM-based pre-training, and medical image-report pre-training counterparts. Codes are available at https://github.com/YtongXie/MedIM.
Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001
Medical Image Anal.3
2024 A deep learning model for the detection of various dementia and MCI pathologies based on resting-state electroencephalography data: A retrospective multicentre study
abstract
Dementia and mild cognitive impairment (MCI) represent significant health challenges in an aging population. As the search for noninvasive, precise and accessible diagnostic methods continues, the efficacy of electroencephalography (EEG) combined with deep convolutional neural networks (DCNNs) in varied clinical settings remains unverified, particularly for pathologies underlying MCI such as Alzheimer's disease (AD), dementia with Lewy bodies (DLB) and idiopathic normal-pressure hydrocephalus (iNPH). Addressing this gap, our study evaluates the generalizability of a DCNN trained on EEG data from a single hospital (Hospital #1). For data from Hospital #1, the DCNN achieved a balanced accuracy (bACC) of 0.927 in classifying individuals as healthy (n = 69) or as having AD, DLB, or iNPH (n = 188). The model demonstrated robustness across institutions, maintaining bACCs of 0.805 for data from Hospital #2 (n = 73) and 0.920 at Hospital #3 (n = 139). Additionally, the model could differentiate AD, DLB, and iNPH cases with bACCs of 0.572 for data from Hospital #1 (n = 188), 0.619 for Hospital #2 (n = 70), and 0.508 for Hospital #3 (n = 139). Notably, it also identified MCI pathologies with a bACC of 0.715 for Hospital #1 (n = 83), despite being trained on overt dementia cases instead of MCI cases. These outcomes confirm the DCNN's adaptability and scalability, representing a significant stride toward its clinical application. Additionally, our findings suggest a potential for identifying shared EEG signatures between MCI and dementia, contributing to the field's understanding of their common pathophysiological mechanisms.
Yusuke Watanabe, Yuki Miyazaki, Masahiro Hata, Ryohei Fukuma, Yasunori Aoki, Hiroaki Kazui, Toshihiko Araki, Daiki Taomoto, Yuto Satake, Takashi Suehiro, Shunsuke Sato, Hideki Kanemoto, Kenji Yoshiyama, Ryouhei Ishii, Tatsuya Harada, Haruhiko Kishima, Manabu Ikeda, Takufumi Yanagisawa
Neural Networks15
2024 Frequency-Aware Feature Fusion for Dense Image Prediction
abstract
Dense image prediction tasks demand features with strong category information and precise spatial boundary details at high resolution. To achieve this, modern hierarchical models often utilize feature fusion, directly adding upsampled coarse features from deep layers and high-resolution features from lower levels. In this paper, we observe rapid variations in fused feature values within objects, resulting in intra-category inconsistency due to disturbed high-frequency features. Additionally, blurred boundaries in fused features lack accurate high frequency, leading to boundary displacement. Building upon these observations, we propose Frequency-Aware Feature Fusion (FreqFusion), integrating an Adaptive Low-Pass Filter (ALPF) generator, an offset generator, and an Adaptive High-Pass Filter (AHPF) generator. The ALPF generator predicts spatially-variant low-pass filters to attenuate high-frequency components within objects, reducing intra-class inconsistency during upsampling. The offset generator refines large inconsistent features and thin boundaries by replacing inconsistent features with more consistent ones through resampling, while the AHPF generator enhances high-frequency detailed boundary information lost during downsampling. Comprehensive visualization and quantitative analysis demonstrate that FreqFusion effectively improves feature consistency and sharpens object boundaries. Extensive experiments across various dense prediction tasks confirm its effectiveness.
Ying Fu 0001, Lin Gu 0003, Chenggang Yan 0001, Tatsuya Harada, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 People Taking Photos That Faces Never Share: Privacy Protection and Fairness Enhancement from Camera to User
abstract
The soaring number of personal mobile devices and public cameras poses a threat to fundamental human rights and ethical principles. For example, the stolen of private information such as face image by malicious third parties will lead to catastrophic consequences. By manipulating appearance of face in the image, most of existing protection algorithms are effective but irreversible. Here, we propose a practical and systematic solution to invertiblely protect face information in the full-process pipeline from camera to final users. Specifically, We design a novel lightweight Flow-based Face Encryption Method (FFEM) on the local embedded system privately connected to the camera, minimizing the risk of eavesdropping during data transmission. FFEM uses a flow-based face encoder to encode each face to a Gaussian distribution and encrypts the encoded face feature by random rotating the Gaussian distribution with the rotation matrix is as the password. While encrypted latent-variable face images are sent to users through public but less reliable channels, password will be protected through more secure channels through technologies such as asymmetric encryption, blockchain, or other sophisticated security schemes. User could select to decode an image with fake faces from the encrypted image on the public channel. Only trusted users are able to recover the original face using the encrypted matrix transmitted in secure channel. More interestingly, by tuning Gaussian ball in latent space, we could control the fairness of the replaced face on attributes such as gender and race. Extensive experiments demonstrate that our solution could protect privacy and enhance fairness with minimal effect on high-level downstream task.
Lin Gu 0003, Sissi Xiaoxiao Wu, Tatsuya Harada, Yingying Zhu 0004
AAAI5
2023 Name Your Colour For the Task: Artificially Discover Colour Naming via Colour Quantisation Transformer
abstract
The long-standing theory that a colour-naming system evolves under dual pressure of efficient communication and perceptual mechanism is supported by more and more linguistic studies, including analysing four decades of diachronic data from the Nafaanra language. This inspires us to explore whether machine learning could evolve and discover a similar colour-naming system via optimising the communication efficiency represented by high-level recognition performance. Here, we propose a novel colour quantisation transformer, CQFormer, that quantises colour space while maintaining the accuracy of machine recognition on the quantised images. Given an RGB image, Annotation Branch maps it into an index map before generating the quantised image with a colour palette; meanwhile the Palette Branch utilises a key-point detection way to find proper colours in the palette among the whole colour space. By interacting with colour annotation, CQFormer is able to balance both the machine vision accuracy and colour perceptual structure such as distinct and stable colour distribution for discovered colour system. Very interestingly, we even observe the consistent evolution pattern between our artificial colour system and basic colour terms across human languages. Besides, our colour quantisation method also offers an efficient quantisation method that effectively compresses the image storage while maintaining high performance in high-level recognition tasks such as classification and detection. Extensive experiments demonstrate the superior performance of our method with extremely low bit-rate colours, showing potential to integrate into quantisation network to quantities from image to network activation. The source code is available at https://github.com/ryeocthiv/CQFormer
Shenghan Su, Lin Gu 0003, Zenghui Zhang, Tatsuya Harada
ICCV5
2023 3D Segmenter: 3D Transformer based Semantic Segmentation via 2D Panoramic Distillation
Zhennan Wu, Yang Li 0193, Yifei Huang 0002, Lin Gu 0003, Tatsuya Harada, Hiroyuki Sato 0002
ICLR5
2023 Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
abstract
To contribute to automating the medical vision-language model, we propose a novel Chest-Xray Different Visual Question Answering (VQA) task. Given a pair of main and reference images, this task attempts to answer several questions on both diseases and, more importantly, the differences between them. This is consistent with the radiologist's diagnosis practice that compares the current image with the reference before concluding the report. We collect a new dataset, namely MIMIC-Diff-VQA, including 700,703 QA pairs from 164,324 pairs of main and reference images. Compared to existing medical VQA datasets, our questions are tailored to the Assessment-Diagnosis-Intervention-Evaluation treatment procedure used by clinical professionals. Meanwhile, we also propose a novel expert knowledge-aware graph representation learning model to address this task. The proposed baseline model leverages expert knowledge such as anatomical structure prior, semantic, and spatial knowledge to construct a multi-relationship graph, representing the image differences between two images for the image difference VQA task. The dataset and code can be found at https://github.com/Holipori/MIMIC-Diff-VQA. We believe this work would further push forward the medical vision language model.
Xinyue Hu 0002, Lin Gu 0003, Qiyuan An, Mengliang Zhang, Kazuma Kobayashi, Tatsuya Harada, Ronald M. Summers, Yingying Zhu 0003
KDD7
2023 Towards AI-Driven Radiology Education: A Self-supervised Segmentation-Based Framework for High-Precision Medical Image Editing
Kazuma Kobayashi, Lin Gu 0003, Ryuichiro Hataya, Mototaka Miyake, Yasuyuki Takamizawa, Sono Ito, Hirokazu Watanabe, Yukihiro Yoshida, Hiroki Yoshimura, Tatsuya Harada, Ryuji Hamamoto
MICCAI (2)10
2023 MedIM: Boost Medical Image Representation via Radiology Report-Guided Masking
Yutong Xie 0001, Lin Gu 0003, Tatsuya Harada, Yong Xia 0001, Qi Wu 0001
MICCAI (1)3
2023 Detection Based Part-level Articulated Object Reconstruction from Single RGBD Image
abstract
We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from previous works that rely on learning instance-level latent space, focusing on man-made articulated objects with predefined part counts. Instead, we propose a novel alternative approach that employs part-level representation, representing instances as combinations of detected parts. While our detect-then-group approach effectively handles instances with diverse part structures and various part counts, it faces issues of false positives, varying part sizes and scales, and an increasing model size due to end-to-end training. To address these challenges, we propose 1) test-time kinematics-aware part fusion to improve detection performance while suppressing false positives, 2) anisotropic scale normalization for part shape learning to accommodate various part sizes and scales, and 3) a balancing strategy for cross-refinement between feature space and output space to improve part detection while maintaining model size. Evaluation on both synthetic and real data demonstrates that our method successfully reconstructs variously structured multiple instances that previous works cannot handle, and outperforms prior works in shape reconstruction and kinematics estimation.
Yuki Kawana, Tatsuya Harada
NeurIPS2
2023 K-VQG: Knowledge-aware Visual Question Generation for Common-sense Acquisition
abstract
Visual Question Generation (VQG) is a task to generate questions from images. When humans ask questions about an image, their goal is often to acquire some new knowledge. However, existing studies on VQG have mainly addressed question generation from answers or question categories, overlooking the objectives of knowledge acquisition. To introduce a knowledge acquisition perspective into VQG, we constructed a novel knowledge-aware VQG dataset called K-VQG. This is the first large, humanly annotated dataset in which questions regarding images are tied to structured knowledge. We also developed a new VQG model that can encode and use knowledge as the target for a question. The experiment results show that our model outperforms existing models on the K-VQG dataset. Our dataset is publicly available at https://uehara-mech.github.io/kvqg.
Kohei Uehara, Tatsuya Harada
WACV2
2023 Backprop Induced Feature Weighting for Adversarial Domain Adaptation with Iterative Label Distribution Alignment
abstract
The requirement for large labeled datasets is one of the limiting factors for training accurate deep neural networks. Unsupervised domain adaptation tackles this problem of limited training data by transferring knowledge from one domain, which has many labeled data, to a different domain for which little to no labeled data is available. One common approach is to learn domain-invariant features for example with an adversarial approach. Previous methods often train the domain classifier and label classifier network separately, where both classification networks have little interaction with each other. In this paper, we introduce a classifier-based backprop-induced weighting of the feature space. This approach has two main advantages. Firstly, it lets the domain classifier focus on features that are important for the classification, and, secondly, it couples the classification and adversarial branch more closely. Furthermore, we introduce an iterative label distribution alignment method, that employs results of previous runs to approximate a class-balanced dataloader. We conduct experiments and ablation studies on three benchmarks Office-31, Office-Home, and DomainNet to show the effectiveness of our proposed algorithm.
Thomas Westfechtel, Hao-Wei Yeh, Meng Qier, Yusuke Mukuta, Tatsuya Harada
WACV5
2023 COMPASS: A creative support system that alerts novelists to the unnoticed missing contents
abstract
Writing a story is never easy. Even experienced writers sometimes unintentionally omit information from their writings, and it makes the stories not understandable from others. Complementing such unintentionally omitted information using a computer is helpful in providing writing support. Recently, in the field of story understanding and generation, story completion (SC) was proposed to generate the missing parts of an incomplete story. Although its applicability is limited because it requires that the user have prior knowledge of the missing part of a story, missing position prediction (MPP) can be used to compensate for this problem. MPP aims to predict the position of the missing part, but the prerequisite knowledge that “one sentence is missing” is still required. In this study, we propose Variable Number MPP (VN-MPP), a new MPP task that removes this restriction; that is, the task to predict multiple missing sentences or to judge whether there are no missing sentences in the first place. We also propose two methods for this new MPP task. One solves this task end-to-end, while the other learns the two modules separately. The latter allows the writer more flexibility in using the information by making the intermediate outputs between the modules explicit. Based on the novel task, we developed a creative writing support system, COMPASS. The results of a user experiment involving professional creators who write texts in Japanese confirm the efficacy and utility of the developed system. This study aimed to propose a creation support system and, at the same time, to build a relationship of trust between creators and researchers to lay the groundwork for future research and development of creation support AI.
Yusuke Mori 0001, Hiroaki Yamane, Ryohei Shimizu, Yusuke Mukuta, Tatsuya Harada
Comput. Speech Lang.5
2023 Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock
Mach. Learn.8
2023 Spherical Image Generation From a Few Normal-Field-of-View Images by Considering Scene Symmetry
abstract
Spherical images taken in all directions (360 degrees by 180 degrees) can represent an entire space including the subject, providing free direction viewing and an immersive experience to viewers. It is convenient and expands the usage scenarios to generate a spherical image from a few normal-field-of-view (NFOV) images, which are partial observations. The primary challenge is generating a plausible image and controlling the high degree of freedom involved in generating a wide area that includes all directions. We focus on scene symmetry, which is a basic property of the global structure of spherical images, such as the rotational and plane symmetries. We propose a method for generating a spherical image from a few NFOV images and controlling the generated regions using scene symmetry. We incorporate the intensity of the symmetry as a latent variable into conditional variational autoencoders to estimate the possible range of symmetry and decode a spherical image whose features are represented through a combination of symmetric transformations of the NFOV image features. Our experiments show that the proposed method can generate various plausible spherical images controlled from asymmetrically to symmetrically, and can reduce the reconstruction errors of the generated images based on the estimated symmetry.
Takayuki Hara, Yusuke Mukuta, Tatsuya Harada
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Correlated and individual feature learning with contrast-enhanced MR for malignancy characterization of hepatocellular carcinoma
abstract
Malignancy characterization of hepatocellular carcinoma (HCC) is of great importance in patient management and prognosis prediction. In this study, we propose an end-to-end correlated and individual feature learning framework to characterize the malignancy of HCC from Contrast-enhanced MR. From the phases of pre-contrast, arterial and portal venous, our framework simultaneously and explicitly learns both the shareable and phase-specific features that are discriminative to malignancy grades. We evaluate our method on the Contrast enhanced MR of 112 consecutive patients with 117 histologically proven HCCs. Experimental results demonstrate that arterial phase yields better results than portal vein and pre-contrast phase. Furthermore, phase specific components show better discriminant ability than the shareable components. Finally, combining the extracted shareable and individual features components has yielded significantly better performance than traditional feature fusion methods. We also conduct t-SNE analysis and feature scoring analysis to qualitatively assess the effectiveness of the proposed method for malignancy characterization.
Yunling Li, Shangxuan Li, Hanqiu Ju, Tatsuya Harada, Honglai Zhang, Ting Duan, Guangyi Wang, Lin Gu 0003, Wu Zhou 0002
Pattern Recognit.4
2022 Fully Spiking Variational Autoencoder
abstract
Spiking neural networks (SNNs) can be run on neuromorphic devices with ultra-high speed and ultra-low energy consumption because of their binary and event-driven nature. Therefore, SNNs are expected to have various applications, including as generative models being running on edge devices to create high-quality images. In this study, we build a variational autoencoder (VAE) with SNN to enable image generation. VAE is known for its stability among generative models; recently, its quality advanced. In vanilla VAE, the latent space is represented as a normal distribution, and floating-point calculations are required in sampling. However, this is not possible in SNNs because all features must be binary time series data. Therefore, we constructed the latent space with an autoregressive SNN model, and randomly selected samples from its output to sample the latent variables. This allows the latent variables to follow the Bernoulli process and allows variational learning. Thus, we build the Fully Spiking Variational Autoencoder where all modules are constructed with SNN. To the best of our knowledge, we are the first to build a VAE only with SNN layers. We experimented with several datasets, and confirmed that it can generate images with the same or better quality compared to conventional ANNs. The code is available at https://github.com/kamata1729/FullySpikingVAE.
Hiromichi Kamata, Yusuke Mukuta, Tatsuya Harada
AAAI3
2022 You Only Need 90K Parameters to Adapt Light: a Light Weight Transformer for Image Enhancement and Exposure Correction
Ziteng Cui, Kunchang Li 0002, Lin Gu 0003, Shenghan Su, Peng Gao 0007, Zhengkai Jiang 0001, Yu Qiao 0001, Tatsuya Harada
BMVC8
2022 Lepard: Learning partial point cloud matching in rigid and deformable scenes
abstract
We present Lepard, a Learning based approach for partial point cloud matching in rigid and deformable scenes. The key characteristics are the following techniques that exploit 3D positional knowledge for point cloud matching: 1) An architecture that disentangles point cloud representation into feature space and 3D position space. 2) A position encoding method that explicitly reveals 3D relative distance information through the dot product of vectors. 3) A repositioning technique that modifies the cross-point-cloud relative positions. Ablation studies demonstrate the effectiveness of the above techniques. In rigid cases, Lepard combined with RANSAC and ICP demonstrates state-of-the-art registration recall of 93.9% / 71.3% on the 3DMatch / 3DLoMatch. In deformable cases, Lepard achieves +27.1% / +34.8% higher non-rigid feature matching recall than the prior art on our newly constructed 4DMatch / 4DLoMatch benchmark. Code and data are available at https://github.com/rabbityl/lepard.
Yang Li 0143, Tatsuya Harada
CVPR2
2022 Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated Objects
abstract
Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure, that is, its joints and how they interact with each other. Unfortunately, assuming the structure to be known, as existing methods do, precludes the ability to work on new object categories. We propose to learn both the appearance and the structure of previously unseen articulated objects by ob-serving them move from multiple views, with no joints annotation supervision, or information about the structure. We observe that 3D points that are static relative to one another should belong to the same part, and that adjacent parts that move relative to each other must be connected by a joint. To leverage this insight, we model the object parts in 3D as ellipsoids, which allows us to identify joints. We combine this explicit representation with an implicit one that compensates for the approximation introduced. We show that our method works for different structures, from quadrupeds, to single-arm robots, to humans. The code is available at https://github.com/NVlabs/watch-it-move and a version of this manuscript that uses animations is at https://arxiv.org/abs/2112.11347
Atsuhiro Noguchi, Umar Iqbal 0001, Jonathan Tremblay, Tatsuya Harada, Orazio Gallo
CVPR4
2022 Revisiting Domain Generalized Stereo Matching Networks from a Feature Consistency Perspective
abstract
Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization capability of stereo matching networks, which has not been adequately considered. Here we address this issue by proposing a simple pixel-wise contrastive learning across the viewpoints. The stereo contrastive feature loss function explicitly constrains the consistency between learned features of matching pixel pairs which are observations of the same 3D points. A stereo selective whitening loss is further introduced to better preserve the stereo feature consistency across domains, which decorrelates stereo features from stereo viewpoint-specific style information. Counter-intuitively, the generalization of feature consistency between two viewpoints in the same scene translates to the generalization of stereo matching performance to unseen domains. Our method is generic in nature as it can be easily embedded into existing stereo networks and does not require access to the samples in the target domain. When trained on synthetic data and generalized to four real-world testing sets, our method achieves superior performance over several state-of-the-art networks. The code is available online11https://github.com/jiaw-z/FCStereo.
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026, Lei Huang 0015, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada, Edwin R. Hancock
CVPR9
2022 Exploring Resolution and Degradation Clues as Self-supervised Signal for Low Quality Object Detection
Ziteng Cui, Yingying Zhu 0004, Lin Gu 0003, Guo-Jun Qi, Renrui Zhang, Zenghui Zhang, Tatsuya Harada
ECCV (9)8
2022 Unsupervised Pose-aware Part Decomposition for Man-Made Articulated Objects
Yuki Kawana, Yusuke Mukuta, Tatsuya Harada
ECCV (3)3
2022 Unsupervised Learning of Efficient Geometry-Aware Neural Articulated Representations
Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada
ECCV (17)4
2022 Deforming Radiance Fields with Cages
Tianhan Xu, Tatsuya Harada
ECCV (33)2
2022 Unsupervised Hierarchical Disentanglement for Video Prediction
abstract
Video prediction is a complicated task as countless possible future frames exist that are equally plausible. While recent work have made progress in the prediction and generation of future video frames, these work have not attempted to disentangle different features of videos such as an object’s structure and its dynamics. Such a disentanglement would allow one to control these aspects to some extent in the prediction phase, while at the same time maintain the object’s intrinsic properties that are learned as the model’s internal representation. In this work, we propose Ladder Variational Recurrent Neural Networks (LVRNN). We employ a type of ladder autoencoder shown to be effective for feature disentanglement on images and apply it to the Variational Recurrent Neural Network (VRNN) architecture, which has been used for video prediction. We rely on extracted keypoints in each frame to separate the structure from the visual features. We then show how different levels of the ladder network learn to disentangle features and demonstrate that each of these levels can be used for controlling different aspects of future frames such as structure and dynamics. We evaluate our method on the Human3.6M and BAIR robot datasets. We show that our method is able to perform hierarchical disentanglement, yet provide reasonable results compared to similar methods.
Mohammad-Reza Motallebi, Thomas Westfechtel, Yang Li 0143, Tatsuya Harada
ICPR4
2022 Boosting Source-free Domain Adaptation via Confidence-based Subsets Feature Alignment
abstract
Source-free Domain Adaptation (SFDA) aims to adapt a model trained on a given (source) environment to the new (target) environment, without directly accessing the source data. Due to the lack of labeled source data, it is often difficult for SFDA methods to provide reliable class representations for the target data. To overcome this issue, we propose the idea of Confidence-based Subsets Feature Alignment (CSFA). CSFA divides the target data into two subsets: confident subset that consists of samples having low entropy class predictions from the source model, and non-confident subset with samples that do not. By using the pseudo-labels from the confident subset, we can frame the original SFDA problem as a Universal Domain Adaptation (UniDA) problem, and provide reliable class representations for the target data by aligning feature distributions of the two subsets. Specifically, we propose a multi-task framework that simultaneously applies a standard SFDA algorithm in combination with a UniDA-inspired algorithm, which further infuses class representations into the adaption process. We evaluate the proposed method on a wide range of cross-domain object recognition tasks and achieve higher or comparable accuracy compared to existing SFDA methods. Ablation studies are conducted to verify the effectiveness of the proposed method.
Hao-Wei Yeh, Thomas Westfechtel, Jia-Bin Huang 0001, Tatsuya Harada
ICPR4
2022 SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate
abstract
The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, emotions, etc. Neural speaker embeddings, trained to identify or verify speakers are typically used to represent and transfer such characteristics from reference speech to synthesized speech.Speech separation on the other hand is the challenging task of separating individual speakers from an overlapping mixed signal of various speakers.Speaker attractors are high-dimensional embedding vectors that pull the time-frequency bins of each speaker's speech towards themselves while repelling those belonging to other speakers.In this work, we explore the possibility of using these powerful speaker attractors for zero-shot speaker adaptation in multi-speaker TTS synthesis and propose speaker attractor text to speech (SATTS).Through various experiments, we show that SATTS can synthesize natural speech from text from an unseen target speaker's reference signal which might have less than ideal recording conditions, i.e. reverberations or mixed with other speakers.
Nabarun Goswami, Tatsuya Harada
INTERSPEECH2
2022 Non-rigid Point Cloud Registration with Neural Deformation Pyramid
abstract
Non-rigid point cloud registration is a key component in many computer vision and computer graphics applications. The high complexity of the unknown non-rigid motion make this task a challenging problem. In this paper, we break down this problem via hierarchical motion decomposition. Our method called Neural Deformation Pyramid (NDP) represents non-rigid motion using a pyramid architecture. Each pyramid level, denoted by a Multi-Layer Perception (MLP), takes as input a sinusoidally encoded 3D point and outputs its motion increments from the previous level. The sinusoidal function starts with a low input frequency and gradually increases when the pyramid level goes down. This allows a multi-level rigid to nonrigid motion decomposition and also speeds up the solving by ×50 times compared to the existing MLP-based approach. Our method achieves advanced partial-to-partial non-rigid point cloud registration results on the 4DMatch/4DLoMatchbenchmark under both no-learned and supervised settings.
Yang Li 0143, Tatsuya Harada
NeurIPS2
2022 Model-Induced Generalization Error Bound for Information-Theoretic Representation Learning in Source-Data-Free Unsupervised Domain Adaptation
abstract
Many unsupervised domain adaptation (UDA) methods have been developed and have achieved promising results in various pattern recognition tasks. However, most existing methods assume that raw source data are available in the target domain when transferring knowledge from the source to the target domain. Due to the emerging regulations on data privacy, the availability of source data cannot be guaranteed when applying UDA methods in a new domain. The lack of source data makes UDA more challenging, and most existing methods are no longer applicable. To handle this issue, this paper analyzes the cross-domain representations in source-data-free unsupervised domain adaptation (SF-UDA). A new theorem is derived to bound the target-domain prediction error using the trained source model instead of the source data. On the basis of the proposed theorem, information bottleneck theory is introduced to minimize the generalization upper bound of the target-domain prediction error, thereby achieving domain adaptation. The minimization is implemented in a variational inference framework using a newly developed latent alignment variational autoencoder (LA-VAE). The experimental results show good performance of the proposed method in several cross-dataset classification tasks without using source data. Ablation studies and feature visualization also validate the effectiveness of our method in SF-UDA.
Baoyao Yang, Hao-Wei Yeh, Tatsuya Harada, Pong C. Yuen
IEEE Trans. Image Process.3
2021 Spherical Image Generation from a Single Image by Considering Scene Symmetry
abstract
Spherical images taken in all directions (360 degrees by 180 degrees) allow the full surroundings of a subject to be represented, providing an immersive experience to viewers. Generating a spherical image from a single normal-field-of-view (NFOV) image is convenient and expands the usage scenarios considerably without relying on a specific panoramic camera or images taken from multiple directions; however, achieving such images remains a challenging and unresolved problem. The primary challenge is controlling the high degree of freedom involved in generating a wide area that includes all directions of the desired spherical image. We focus on scene symmetry, which is a basic property of the global structure of spherical images, such as rotational symmetry, plane symmetry, and asymmetry. We propose a method for generating a spherical image from a single NFOV image and controlling the degree of freedom of the generated regions using the scene symmetry. To estimate and control the scene symmetry using both a circular shift and flip of the latent image features, we incorporate the intensity of the symmetry as a latent variable into conditional variational autoencoders. Our experiments show that the proposed method can generate various plausible spherical images controlled from symmetric to asymmetric, and can reduce the reconstruction errors of the generated images based on the estimated symmetry.
Takayuki Hara, Yusuke Mukuta, Tatsuya Harada
AAAI3
2021 Leveraging Human Selective Attention for Medical Image Analysis with Limited Training Data
Yifei Huang 0002, Lijin Yang, Lin Gu 0003, Yingying Zhu 0004, Hirofumi Seo, Qiuming Meng, Tatsuya Harada, Yoichi Sato 0001
BMVC8
2021 Blur, Noise, and Compression Robust Generative Adversarial Networks
abstract
Generative adversarial networks (GANs) have gained considerable attention owing to their ability to reproduce images. However, they can recreate training images faithfully despite image degradation in the form of blur, noise, and compression, generating similarly degraded images. To solve this problem, the recently proposed noise robust GAN (NR-GAN) provides a partial solution by demonstrating the ability to learn a clean image generator directly from noisy images using a two-generator model comprising image and noise generators. However, its application is limited to noise, which is relatively easy to decompose owing to its additive and reversible characteristics, and its application to irreversible image degradation, in the form of blur, compression, and combination of all, remains a challenge. To address these problems, we propose blur, noise, and compression robust GAN (BNCR-GAN) that can learn a clean image generator directly from degraded images without knowledge of degradation parameters (e.g., blur kernel types, noise amounts, or quality factor values). Inspired by NR-GAN, BNCR-GAN uses a multiple-generator model composed of image, blur-kernel, noise, and quality-factor generators. However, in contrast to NR-GAN, to address irreversible characteristics, we introduce masking architectures adjusting degradation strength values in a data-driven manner using bypasses before and after degradation. Furthermore, to suppress uncertainty caused by the combination of blur, noise, and compression, we introduce adaptive consistency losses imposing consistency between irreversible degradation processes according to the degradation strengths. We demonstrate the effectiveness of BNCR-GAN through large-scale comparative studies on CIFAR-10 and a generality analysis on FFHQ. In addition, we demonstrate the applicability of BNCR-GAN in image restoration.
Takuhiro Kaneko, Tatsuya Harada
CVPR2
2021 Goal-Oriented Gaze Estimation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial for visual-semantic embedding. Interestingly, when recognizing unseen images, human would also automatically gaze at regions with certain semantic clue. Therefore, we introduce a novel goal-oriented gaze estimation module (GEM) to improve the discriminative attribute localization based on the class-level attributes for ZSL. We aim to predict the actual human gaze location to get the visual attention regions for recognizing a novel object guided by attribute description. Specifically, the task-dependent attention is learned with the goal-oriented GEM, and the global image features are simultaneously optimized with the regression of local attribute features. Experiments on three ZSL benchmarks, i.e., CUB, SUN and AWA2, show the superiority or competitiveness of our proposed method against the state-of-the-art ZSL methods. The ablation analysis on real gaze data CUB-VWSW also validates the benefits and accuracy of our gaze estimation module. This work implies the promising benefits of collecting human gaze dataset and automatic gaze estimation algorithms on high-level computer vision tasks. The code is available at https://github.com/osierboy/GEM-ZSL.
Yang Liu 0357, Lei Zhou 0008, Xiao Bai 0001, Yifei Huang 0002, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada
CVPR7
2021 Multitask AET with Orthogonal Tangent Regularity for Dark Object Detection
abstract
Dark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET) model which is able to explore the intrinsic pattern behind illumination translation. In a self-supervision manner, the MAET learns the intrinsic visual structure by encoding and decoding the realistic illumination-degrading transformation considering the physical noise model and image signal processing (ISP). Based on this representation, we achieve the object detection task by decoding the bounding box coordinates and classes. To avoid the over-entanglement of two tasks, our MAET disentangles the object and degrading features by imposing an orthogonal tangent regularity. This forms a parametric manifold along which multitask predictions can be geometrically formulated by maximizing the orthogonality between the tangents along the outputs of respective tasks. Our framework can be implemented based on the mainstream object detection architecture and directly trained end-to-end using normal target detection datasets, such as VOC and COCO. We have achieved the state-of-the-art performance using synthetic and real-world datasets. Codes will be released at https://github.com/cuiziteng/MAET.
Ziteng Cui, Guo-Jun Qi, Lin Gu 0003, Shaodi You, Zenghui Zhang, Tatsuya Harada
ICCV6
2021 Neural Articulated Radiance Field
abstract
We present Neural Articulated Radiance Field (NARF), a novel deformable 3D representation for articulated objects learned from images. While recent advances in 3D implicit representation have made it possible to learn models of complex objects, learning pose-controllable representations of articulated objects remains a challenge, as current methods require 3D shape supervision and are unable to render appearance. In formulating an implicit representation of 3D articulated objects, our method considers only the rigid transformation of the most relevant object part in solving for the radiance field at each 3D location. In this way, the proposed method represents pose-dependent changes without significantly increasing the computational complexity. NARF is fully differentiable and can be trained from images with pose annotations. Moreover, through the use of an autoencoder, it can learn appearance variations over multiple instances of an object class. Experiments show that the proposed method is efficient and can generalize well to novel poses. The code is available for research purposes at https://github.com/nogu-atsu/NARF.
Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada
ICCV4
2021 Hyperbolic Neural Networks++
Ryohei Shimizu, Yusuke Mukuta, Tatsuya Harada
ICLR3
2021 Real-Time Mesh Extraction from Implicit Functions via Direct Reconstruction of Decision Boundary
abstract
The ability to estimate 3D object shape from a single image is vital to robotics and manufacturing. For instance, it enables iterative trial-and-error in simulated environments. In single-view reconstruction, implicit functions have demonstrated superior results over traditional methods. However, implicit functions suffer from the heavy computation of mesh extraction. This is due to the indirect mesh extraction, where the number of evaluation points grows cubically with resolution. On the other hand, reducing the resolution results in the discretization error of marching cubes (MC). In this work, we aim to perform efficient and accurate mesh extraction from implicit functions. The idea is to directly reconstruct the decision boundary of implicit functions as a mesh by reverse tracing from the output. It eliminates the need for evaluating massive points and error-prone MC. Consequently, we propose implementing an implicit function via a composite function of a flow and Binary-coded Input Neural Network (BCINN). The boundary of BCINN is easily identifiable, and the flow is invertible. Owing to these properties, the decision boundary of the composite function can be directly and efficiently reconstructed. In our experiments, we demonstrate that the proposed method significantly improves runtime/memory efficiency, with results comparable to those of existing methods. Specifically, our method enables real-time high-quality mesh inference from a single image.
Wataru Kawai, Yusuke Mukuta, Tatsuya Harada
ICRA3
2021 Making Video Recognition Models Robust to Common Corruptions With Supervised Contrastive Learning
abstract
The video understanding capability of video recognition models has been significantly improved by the development of deep learning techniques and various video datasets available. However, video recognition models are still vulnerable to invisible perturbations, which limits the use of deep video recognition models in the real world. We present a new benchmark for the robustness of action recognition classifiers to general corruptions, and show that a supervised contrastive learning framework is effective in obtaining discriminative and stable video representations, and makes deep video recognition models robust to general input corruptions. Experiments on the action recognition task for corrupted videos show the high robustness of the proposed method on the UCF101 and HMDB51 datasets with various common corruptions.
Tomu Hirata, Yusuke Mukuta, Tatsuya Harada
MMAsia3
2021 Generation of Variable-Length Time Series from Text using Dynamic Time Warping-Based Method
abstract
This study is aimed at finding a suitable method for generating time-series data such as video clips or avatar motions from text stating multiple events. This paper addresses the generation of variable-length time-series data considering the order and variable duration of events stated in the text. Although the use of the variant of Mean Squared Error (MSE) is a common means of training, only the gap between the element of ground-truth (GT) data and generated data at the same time are considered. Thus, variants of MSE are unsuitable for the task at hand because the loss may not be small for the generated and GT data with the same order of events if the time for each event does not overlap. To solve the problem, we propose a Dynamic Time Warping-Like method for Variable-Length data (DTWL-VL), which determines the corresponding elements of the GT and the generated data, allowing for the time difference between them, and makes them closer. We compared DTWL-VL, a variant of MSE, and an existing method for time-series data generation which considers the time difference between the corresponding part in the GT and generated data. Since the existing method is aimed at generating fixed-length data, we extend the method for generating variable-length time-series data. We conducted experiments using a dataset prepared for this study. Both DTWL-VL and the existing methods outperformed the MSE variant. Moreover, although the existing method outperformed DTWL-VL under certain settings, DTWL-VL required a smaller training period.
Ayaka Ideno, Yusuke Mukuta, Tatsuya Harada
MMAsia3
2021 SoFA: Source-data-free Feature Alignment for Unsupervised Domain Adaptation
abstract
Applying a trained model on a new scenario may suffer from domain shift. Unsupervised domain adaptation (UDA) has been proven to be an effective approach to solve the problem of domain shift by leveraging both data from the scenario that the model was trained on (source) and the new scenario (target). Although the source data are available for training the source model, there is no guarantee that the source data will still be available when applying UDA in the future due to emerging regulations on privacy of data. This results in the in-applicability of most existing UDA methods in the absence of source data. This paper proposes a source-data-free feature alignment (SoFA) method to address this problem by only using the trained source model and unlabeled target data. The source model is used to predict the labels for target data, and we model the generation process from predicted classes to input data to infer the latent features for alignment. Specifically, a mixture of Gaussian distributions is induced from the predicted classes as the reference distribution. The encoded target features are then aligned to the reference distribution via variational inference to extract class semantics without accessing source data. Relationship of the proposed method and the theory of domain adaptation is provided to verify the performance. Experimental results show the proposed method achieves higher or comparable accuracy compared to the existing methods in several cross-dataset classification tasks. Ablation studies are also conducted to confirm the importance of latent feature alignment to adaptation performance.
Hao-Wei Yeh, Baoyao Yang, Pong C. Yuen, Tatsuya Harada
WACV4
2021 Humor meets morality: Joke generation based on moral judgement
abstract
Although humor enriches human lives, some jokes fail to amuse people because of a lack of morality. In this paper, we propose a mechanism capable of selecting humor based on moral criteria. To this end, we first construct a model based on an N -gram corpus and generate joke candidates using various template patterns. We then employ a moral judgement classifier based on a recurrent neural network and utilize the trained model for humor selection. The experimental results obtained from best–worst scaling demonstrate that this scheme is able to generate jokes with moral category labels. We confirmed that jokes about the classifier categorized as Loyalty and Authority , which are regarded as good in our study, are funnier than jokes about Fairness , Purity , Harm , Cheating , and Degradation . Although we did not confirm that there was a difference in the funny level between good and bad moral jokes, the results demonstrate that moral categories of humor can affect the funny level.
Hiroaki Yamane, Yusuke Mori 0001, Tatsuya Harada
Inf. Process. Manag.3
2021 Decomposing normal and abnormal features of medical images for content-based image retrieval of glioma imaging
abstract
In medical imaging, the characteristics purely derived from a disease should reflect the extent to which abnormal findings deviate from the normal features. Indeed, physicians often need corresponding images without abnormal findings of interest or, conversely, images that contain similar abnormal findings regardless of normal anatomical context. This is called comparative diagnostic reading of medical images, which is essential for a correct diagnosis. To support comparative diagnostic reading, content-based image retrieval (CBIR) that can selectively utilize normal and abnormal features in medical images as two separable semantic components will be useful. In this study, we propose a neural network architecture to decompose the semantic components of medical images into two latent codes: normal anatomy code and abnormal anatomy code. The normal anatomy code represents counterfactual normal anatomies that should have existed if the sample is healthy, whereas the abnormal anatomy code attributes to abnormal changes that reflect deviation from the normal baseline. By calculating the similarity based on either normal or abnormal anatomy codes or the combination of the two codes, our algorithm can retrieve images according to the selected semantic component from a dataset consisting of brain magnetic resonance images of gliomas. Moreover, it can utilize a synthetic query vector combining normal and abnormal anatomy codes from two different query images. To evaluate whether the retrieved images are acquired according to the targeted semantic component, the overlap of the ground-truth labels is calculated as metrics of the semantic consistency. Our algorithm provides a flexible CBIR framework by handling the decomposed features with qualitatively and quantitatively remarkable results.
Kazuma Kobayashi, Ryuichiro Hataya, Yusuke Kurose, Mototaka Miyake, Masamichi Takahashi, Akiko Nakagawa, Tatsuya Harada, Ryuji Hamamoto
Medical Image Anal.7
2021 View-invariant action recognition via Unsupervised AttentioN Transfer (UANT)
Yanli Ji, Yang Yang 0002, Heng Tao Shen, Tatsuya Harada
Pattern Recognit.4
2020 Domain Generalization Using a Mixture of Multiple Latent Domains
abstract
When domains, which represent underlying data distributions, vary during training and testing processes, deep neural networks suffer a drop in their performance. Domain generalization allows improvements in the generalization performance for unseen target domains by using multiple source domains. Conventional methods assume that the domain to which each sample belongs is known in training. However, many datasets, such as those collected via web crawling, contain a mixture of multiple latent domains, in which the domain of each sample is unknown. This paper introduces domain generalization using a mixture of multiple latent domains as a novel and more realistic scenario, where we try to train a domain-generalized model without using domain labels. To address this scenario, we propose a method that iteratively divides samples into latent domains via clustering, and which trains the domain-invariant feature extractor shared among the divided latent domains via adversarial learning. We assume that the latent domain of images is reflected in their style, and thus, utilize style features for clustering. By using these features, our proposed method successfully discovers latent domains and achieves domain generalization even if the domain labels are not given. Experiments show that our proposed method can train a domain-generalized model without using domain labels. Moreover, it outperforms conventional domain generalization methods, including those that utilize domain labels.
Toshihiko Matsuura, Tatsuya Harada
AAAI2
2020 Accurate Parts Visualization for Explaining CNN Reasoning via Semantic Segmentation
Ren Harada, Antonio Tejero-de-Pablos, Tatsuya Harada
BMVC3
2020 Noise Robust Generative Adversarial Networks
abstract
Generative adversarial networks (GANs) are neural networks that learn data distributions through adversarial training. In intensive studies, recent GANs have shown promising results for reproducing training images. However, in spite of noise, they reproduce images with fidelity. As an alternative, we propose a novel family of GANs called noise robust GANs (NR-GANs), which can learn a clean image generator even when training images are noisy. In particular, NR-GANs can solve this problem without having complete noise information (e.g., the noise distribution type, noise amount, or signal-noise relationship). To achieve this, we introduce a noise generator and train it along with a clean image generator. However, without any constraints, there is no incentive to generate an image and noise separately. Therefore, we propose distribution and transformation constraints that encourage the noise generator to capture only the noise-specific components. In particular, considering such constraints under different assumptions, we devise two variants of NR-GANs for signal-independent noise and three variants of NR-GANs for signal-dependent noise. On three benchmark datasets, we demonstrate the effectiveness of NR-GANs in noise robust image generation. Furthermore, we show the applicability of NR-GANs in image denoising. Our code is available at https://github.com/takuhirok/NR-GAN/.
Takuhiro Kaneko, Tatsuya Harada
CVPR2
2020 Learning to Optimize Non-Rigid Tracking
abstract
One of the widespread solutions for non-rigid tracking has a nested-loop structure: with Gauss-Newton to minimize a tracking objective in the outer loop, and Preconditioned Conjugate Gradient (PCG) to solve a sparse linear system in the inner loop. In this paper, we employ learnable optimizations to improve tracking robustness and speed up solver convergence. First, we upgrade the tracking objective by integrating an alignment data term on deep features which are learned end-to-end through CNN. The new tracking objective can capture the global deformation which helps Gauss-Newton to jump over local minimum, leading to robust tracking on large non-rigid motions. Second, we bridge the gap between the preconditioning technique and learning method by introducing a ConditionNet which is trained to generate a preconditioner such that PCG can converge within a small number of steps. Experimental results indicate that the proposed learning method converges faster than the original PCG by a large margin.
Yang Li 0143, Aljaz Bozic, Tianwei Zhang 0002, Yanli Ji, Tatsuya Harada, Matthias Nießner
CVPR5
2020 Bounding-Box Channels for Visual Relationship Detection
Sho Inayoshi, Keita Otani, Antonio Tejero-de-Pablos, Tatsuya Harada
ECCV (5)4
2020 RGBD-GAN: Unsupervised 3D Representation Learning From Natural Image Datasets via RGBD Image Synthesis
Atsuhiro Noguchi, Tatsuya Harada
ICLR2
2020 SplitFusion: Simultaneous Tracking and Mapping for Non-Rigid Scenes
abstract
We present SplitFusion, a novel dense RGB-D SLAM framework that simultaneously performs tracking and dense reconstruction for both rigid and non-rigid components of the scene. SplitFusion first adopts deep learning based semantic instant segmentation technique to split the scene into rigid or non-rigid surfaces. The split surfaces are independently tracked via rigid or non-rigid ICP and reconstructed through incremental depth map fusion. Experimental results show that the proposed approach can provide not only accurate environment maps but also well-reconstructed non-rigid targets, e.g., the moving humans.
Yang Li 0143, Tianwei Zhang 0002, Yoshihiko Nakamura, Tatsuya Harada
IROS4
2020 Point Cloud Based Reinforcement Learning for Sim-to-Real and Partial Observability in Visual Navigation
abstract
Reinforcement Learning (RL), among other learning-based methods, represents powerful tools to solve complex robotic tasks (e.g., actuation, manipulation, navigation, etc.), with the need for real-world data to train these systems as one of its most important limitations. The use of simulators is one way to address this issue, yet knowledge acquired in simulations does not work directly in the real-world, which is known as the sim-to-real transfer problem. While previous works focus on the nature of the images used as observations (e.g., textures and lighting), which has proven useful for a sim-to-sim transfer, they neglect other concerns regarding said observations, such as precise geometrical meanings, failing at robot-to-robot, and thus in sim-to-real transfers. We propose a method that learns on an observation space constructed by point clouds and environment randomization, generalizing among robots and simulators to achieve sim-to-real, while also addressing partial observability. We demonstrate the benefits of our methodology on the point goal navigation task, in which our method proves to be highly unaffected to unseen scenarios produced by robot-to-robot transfer, outperforms image-based baselines in robot-randomized experiments, and presents high performances in sim-to-sim conditions. Finally, we perform several experiments to validate the sim-to-real transfer to a physical domestic robot platform, confirming the out-of-the-box performance of our system.
Kenzo Lobos-Tsunekawa, Tatsuya Harada
IROS2
2020 Learning Agile Locomotion via Adversarial Training
abstract
Developing controllers for agile locomotion is a long-standing challenge for legged robots. Reinforcement learning (RL) and Evolution Strategy (ES) hold the promise of automating the design process of such controllers. However, dedicated and careful human effort is required to design training environments to promote agility. In this paper, we present a multi-agent learning system, in which a quadruped robot (protagonist) learns to chase another robot (adversary) while the latter learns to escape. We find that this adversarial training process not only encourages agile behaviors but also effectively alleviates the laborious environment design effort. In contrast to prior works that used only one adversary, we find that training an ensemble of adversaries, each of which specializes in a different escaping strategy, is essential for the protagonist to master agility. Through extensive experiments, we show that the locomotion controller learned with adversarial training significantly outperforms carefully designed baselines.
Yujin Tang, Jie Tan 0001, Tatsuya Harada
IROS3
2020 Neural Star Domain as Primitive Representation
abstract
Reconstructing 3D objects from 2D images is a fundamental task in computer vision. Acurate structured reconstruction by parsimonious and semantic primitive representation further broadens its application. When reconstructing a target shape with multiple primitives, it is preferable that one can instantly access the union of basic properties of the shape such as collective volume and surface, treating the primitives as if they are one single shape. This becomes possible by primitive representation with unified implicit and explicit representations. However, primitive representations in current approaches do not satisfy all of the above requirements at the same time. To solve this problem, we propose a novel primitive representation named neural star domain (NSD) that learns primitive shapes in the star domain. We show that NSD is a universal approximator of the star domain and is not only parsimonious and semantic but also an implicit and explicit shape representation. We demonstrate that our approach outperforms existing methods in image reconstruction tasks, semantic capabilities, and speed and quality of sampling high-resolution meshes.
Yuki Kawana, Yusuke Mukuta, Tatsuya Harada
NeurIPS3
2019 Estimating the Causal Effect from Partially Observed Time Series
abstract
Many real-world systems involve interacting time series. The ability to detect causal dependencies between system components from observed time series of their outputs is essential for understanding system behavior. The quantification of causal influences between time series is based on the definition of some causality measure. Partial Canonical Correlation Analysis (Partial CCA) and its extensions are examples of methods used for robustly estimating the causal relationships between two multidimensional time series even when the time series are short. These methods assume that the input data are complete and have no missing values. However, real-world data often contain missing values. It is therefore crucial to estimate the causality measure robustly even when the input time series is incomplete. Treating this problem as a semi-supervised learning problem, we propose a novel semi-supervised extension of probabilistic Partial CCA called semi-Bayesian Partial CCA. Our method exploits the information in samples with missing values to prevent the overfitting of parameter estimation even when there are few complete samples. Experiments based on synthesized and real data demonstrate the ability of the proposed method to estimate causal relationships more correctly than existing methods when the data contain missing values, the dimensionality is large, and the number of samples is small.
Akane Iseki, Yusuke Mukuta, Yoshitaka Ushiku, Tatsuya Harada
AAAI4
2019 Class-Distinct and Class-Mutual Image Generation with GANs
Takuhiro Kaneko, Yoshitaka Ushiku, Tatsuya Harada
BMVC3
2019 Learning to Explain With Complemental Examples
abstract
This paper addresses the generation of explanations with visual examples. Given an input sample, we build a system that not only classifies it to a specific category, but also outputs linguistic explanations and a set of visual examples that render the decision interpretable. Focusing especially on the complementarity of the multimodal information, i.e., linguistic and visual examples, we attempt to achieve it by maximizing the interaction information, which provides a natural definition of complementarity from an information theoretical viewpoint. We propose a novel framework to generate complemental explanations, on which the joint distribution of the variables to explain, and those to be explained is parameterized by three different neural networks: predictor, linguistic explainer, and example selector. Explanation models are trained collaboratively to maximize the interaction information to ensure the generated explanation are complemental to each other for the target. The results of experiments conducted on several datasets demonstrate the effectiveness of the proposed method.
Atsushi Kanehira, Tatsuya Harada
CVPR2
2019 Multimodal Explanations by Predicting Counterfactuality in Videos
abstract
This study addresses generating counterfactual explanations with multimodal information. Our goal is not only to classify a video into a specific category, but also to provide explanations on why it is not categorized to a specific class with combinations of visual-linguistic information. Requirements that the expected output should satisfy are referred to as counterfactuality in this paper: (1) Compatibility of visual-linguistic explanations, and (2) Positiveness/negativeness for the specific positive/negative class. Exploiting a spatio-temporal region (tube) and an attribute as visual and linguistic explanations respectively, the explanation model is trained to predict the counterfactuality for possible combinations of multimodal information in a post-hoc manner. The optimization problem, which appears during training/inference, can be efficiently solved by inserting a novel neural network layer, namely the maximum subpath layer. We demonstrated the effectiveness of this method by comparison with a baseline of the action recognition datasets extended for this task. Moreover, we provide information-theoretical insight into the proposed method.
Atsushi Kanehira, Kentaro Takemoto, Sho Inayoshi, Tatsuya Harada
CVPR4
2019 Label-Noise Robust Generative Adversarial Networks
abstract
Generative adversarial networks (GANs) are a framework that learns a generative distribution through adversarial training. Recently, their class conditional extensions (e.g., conditional GAN (cGAN) and auxiliary classifier GAN (AC-GAN)) have attracted much attention owing to their ability to learn the disentangled representations and to improve the training stability. However, their training requires the availability of large-scale accurate class-labeled data, which are often laborious or impractical to collect in a real-world scenario. To remedy this, we propose a novel family of GANs called label-noise robust GANs (rGANs), which, by incorporating a noise transition model, can learn a clean label conditional generative distribution even when training labels are noisy. In particular, we propose two variants: rAC-GAN, which is a bridging model between AC-GAN and the label-noise robust classification model, and rcGAN, which is an extension of cGAN and solves this problem with no reliance on any classifier. In addition to providing the theoretical background, we demonstrate the effectiveness of our models through extensive experiments using diverse GAN configurations, various noise settings, and multiple evaluation metrics (in which we tested 402 conditions in total).
Takuhiro Kaneko, Yoshitaka Ushiku, Tatsuya Harada
CVPR3
2019 Learning View Priors for Single-View 3D Reconstruction
abstract
There is some ambiguity in the 3D shape of an object when the number of observed views is small. Because of this ambiguity, although a 3D object reconstructor can be trained using a single view or a few views per object, reconstructed shapes only fit the observed views and appear incorrect from the unobserved viewpoints. To reconstruct shapes that look reasonable from any viewpoint, we propose to train a discriminator that learns prior knowledge regarding possible views. The discriminator is trained to distinguish the reconstructed views of the observed viewpoints from those of the unobserved viewpoints. The reconstructor is trained to correct unobserved views by fooling the discriminator. Our method outperforms current state-of-the-art methods on both synthetic and natural image datasets; this validates the effectiveness of our method.
Hiroharu Kato, Tatsuya Harada
CVPR2
2019 Strong-Weak Distribution Alignment for Adaptive Object Detection
abstract
We propose an approach for unsupervised adaptation of object detectors from label-rich to label-poor domains which can significantly reduce annotation costs associated with detection. Recently, approaches that align distributions of source and target images using an adversarial loss have been proven effective for adapting object classifiers. However, for object detection, fully matching the entire distributions of source and target images to each other at the global image level may fail, as domains could have distinct scene layouts and different combinations of objects. On the other hand, strong matching of local features such as texture and color makes sense, as it does not change category level semantics. This motivates us to propose a novel method for detector adaptation based on strong local alignment and weak global alignment. Our key contribution is the weak alignment model, which focuses the adversarial alignment loss on images that are globally similar and puts less emphasis on aligning images that are globally dissimilar. Additionally, we design the strong domain alignment model to only look at local receptive fields of the feature map. We empirically verify the effectiveness of our method on four datasets comprising both large and small domain shifts. Our code is available at https://github.com/VisionLearningGroup/DA_Detection.
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, Kate Saenko
CVPR3
2019 Image Generation From Small Datasets via Batch Statistics Adaptation
abstract
Thanks to the recent development of deep generative models, it is becoming possible to generate high-quality images with both fidelity and diversity. However, the training of such generative models requires a large dataset. To reduce the amount of data required, we propose a new method for transferring prior knowledge of the pre-trained generator, which is trained with a large dataset, to a small dataset in a different domain. Using such prior knowledge, the model can generate images leveraging some common sense that cannot be acquired from a small dataset. In this work, we propose a novel method focusing on the parameters for batch statistics, scale and shift, of the hidden layers in the generator. By training only these parameters in a supervised manner, we achieved stable training of the generator, and our method can generate higher quality images compared to previous methods without collapsing, even when the dataset is small (~100). Our results show that the diversity of the filters acquired in the pre-trained generator is important for the performance on the target domain. Our method makes it possible to add a new class or domain to a pre-trained generator without disturbing the performance on the original domain. Code is available at github.com/nogu-atsu/small-dataset-image-generation.
Atsuhiro Noguchi, Tatsuya Harada
ICCV2
2019 Multi-Stage Pathological Image Classification Using Semantic Segmentation
abstract
Histopathological image analysis is an essential process for the discovery of diseases such as cancer. However, it is challenging to train CNN on whole slide images (WSIs) of gigapixel resolution considering the available memory capacity. Most of the previous works divide high resolution WSIs into small image patches and separately input them into the model to classify it as a tumor or a normal tissue. However, patch-based classification uses only patch-scale local information but ignores the relationship between neighboring patches. If we consider the relationship of neighboring patches and global features, we can improve the classification performance. In this paper, we propose a new model structure combining the patch-based classification model and whole slide-scale segmentation model in order to improve the prediction performance of automatic pathological diagnosis. We extract patch features from the classification model and input them into the segmentation model to obtain a whole slide tumor probability heatmap. The classification model considers patch-scale local features, and the segmentation model can take global information into account. We also propose a new optimization method that retains gradient information and trains the model partially for end-to-end learning with limited GPU memory capacity. We apply our method to the tumor/normal prediction on WSIs and the classification performance is improved compared with the conventional patch-based method.
Shusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe, Masashi Fukayama, Akihiko Yoshizawa, Masanobu Kitagawa, Tatsuya Harada
ICCV8
2019 Generating Easy-to-Understand Referring Expressions for Target Identifications
abstract
This paper addresses the generation of referring expressions that not only refer to objects correctly but also let humans find them quickly. As a target becomes relatively less salient, identifying referred objects itself becomes more difficult. However, the existing studies regarded all sentences that refer to objects correctly as equally good, ignoring whether they are easily understood by humans. If the target is not salient, humans utilize relationships with the salient contexts around it to help listeners to comprehend it better. To derive this information from human annotations, our model is designed to extract information from the target and from the environment. Moreover, we regard that sentences that are easily understood are those that are comprehended correctly and quickly by humans. We optimized this by using the time required to locate the referred objects by humans and their accuracies. To evaluate our system, we created a new referring expression dataset whose images were acquired from Grand Theft Auto V (GTA V), limiting targets to persons. Experimental results show the effectiveness of our approach. Our code and dataset are available at https://github.com/mikittt/easy-to-understand-REG.
Mikihiro Tanaka, Takayuki Itamochi, Kenichi Narioka, Ikuro Sato, Yoshitaka Ushiku, Tatsuya Harada
ICCV6
2019 Improved Optical Flow for Gesture-based Human-robot Interaction
abstract
Gesture interaction is a natural way of communicating with a robot as an alternative to speech. Gesture recognition methods leverage optical flow in order to understand human motion. However, while accurate optical flow estimation (i.e., traditional) methods are costly in terms of runtime, fast estimation (i.e., deep learning) methods' accuracy can be improved. In this paper, we present a pipeline for gesture-based human-robot interaction that uses a novel optical flow estimation method in order to achieve an improved speed-accuracy trade-off. Our optical flow estimation method introduces four improvements to previous deep learning-based methods: strong feature extractors, attention to contours, midway features, and a combination of these three. This results in a better understanding of motion, and a finer representation of silhouettes. In order to evaluate our pipeline, we generated our own dataset, MIBURI, which contains gestures to command a house service robot. In our experiments, we show how our method improves not only optical flow estimation, but also gesture recognition, offering a speed-accuracy trade-off more realistic for practical robot applications.
Jen-Yen Chang, Antonio Tejero-de-Pablos, Tatsuya Harada
ICRA3
2019 Pose Graph optimization for Unsupervised Monocular Visual Odometry
abstract
Unsupervised Learning based monocular visual odometry (VO) has lately drawn significant attention for its potential in label-free leaning ability and robustness to camera parameters and environmental variations. However, partially due to the lack of drift correction technique, these methods are still by far less accurate than geometric approaches for large-scale odometry estimation. In this paper, we propose to leverage graph optimization and loop closure detection to overcome limitations of unsupervised learning based monocular visual odometry. To this end, we propose a hybrid VO system which combines an unsupervised monocular VO called NeuralBundler with a pose graph optimization back-end. NeuralBundler is a neural network architecture that uses temporal and spatial photometric loss as main supervision and generates a windowed pose graph consists of multi-view 6DoF constraints. We propose a novel pose cycle consistency loss to relieve the tensions in the windowed pose graph, leading to improved performance and robustness. In the back-end, a global pose graph is built from local and loop 6DoF constraints estimated by NeuralBundler, and is optimized over SE(3). Empirical evaluation on the KITTI odometry dataset demonstrates that 1) NeuralBundler achieves state-of-the-art performance on unsupervised monocular VO estimation, and 2) our whole approach can achieve efficient loop closing and show favorable overall translational accuracy compared to established monocular SLAM systems.
Yang Li 0143, Yoshitaka Ushiku, Tatsuya Harada
ICRA3
2019 Detecting, Opening and Navigating through Doors: A Unified Framework for Human Service Robots
Francesco Savarese, Antonio Tejero-de-Pablos, Stefano Quer, Tatsuya Harada
ICSOFT4
2019 Toward a Better Story End: Collecting Human Evaluation with Reasons
abstract
Creativity is an essential element of human nature used for many activities, such as telling a story.Based on human creativity, researchers have attempted to teach a computer to generate stories automatically or support this creative process.In this study, we undertake the task of story ending generation.This is a relatively new task, in which the last sentence of a given incomplete story is automatically generated.This is challenging because, in order to predict an appropriate ending, the generation method should comprehend the context of events.Despite the importance of this task, no clear evaluation metric has been established thus far; hence, it has remained an open problem.Therefore, we study the various elements involved in evaluating an automatic method for generating story endings.First, we introduce a baseline hierarchical sequenceto-sequence method for story ending generation.Then, we conduct a pairwise comparison against human-written endings, in which annotators choose the preferable ending.In addition to a quantitative evaluation, we conduct a qualitative evaluation by asking annotators to specify the reason for their choice.From the collected reasons, we discuss what elements the evaluation should focus on, to thereby propose effective metrics for the task.
Yusuke Mori 0001, Hiroaki Yamane, Yusuke Mukuta, Tatsuya Harada
INLG4
2019 Simultaneous Transparent and Non-Transparent Object Segmentation With Multispectral Scenes
abstract
For an autonomous mobile system such as an autonomous robot that moves throughout a city, semantic segmentation is important. Performing semantic segmentation under diverse conditions, in turn, requires 1) a robust ability to recognize objects in low-visibility environments, such as at night and 2) the ability to recognize objects that transmit visible light, such as glass and acrylic used in doors and windows. To satisfy these requirements, using RGB images and infrared images simultaneously is considered effective. Visibility and infrared transmission characteristics are different for different objects; therefore, merely entering them into the conventional semantic segmentation framework is not applicable. For example, when a pedestrian is present behind a glass, the visible image captures the pedestrian rather than the glass and the infrared image captures the glass. In this research, we propose a new semantic segmentation method having a three-stream structure, focusing on the difference in the transmission characteristics. This method extracts not only valid features for ordinary non-transparent objects but also features effective for the recognition of transparent objects by utilizing differences in objects to be imaged owing to transmission characteristics. Furthermore, we constructed a new dataset called “coaxials” for the visible and infrared coaxial dataset and demonstrated that we can obtain better segmentation performance compared with the conventional method.
Atsuro Okazawa, Tomoyuki Takahata, Tatsuya Harada
IROS3
2019 Gastric Cancer Detection from Endoscopic Images Using Synthesis by GAN
Teppei Kanayama, Yusuke Kurose, Kiyohito Tanaka, Kento Aida, Shin'ichi Satoh 0001, Masaru Kitsuregawa, Tatsuya Harada
MICCAI (5)7
2019 Texture-Based Classification of Significant Stenosis in CCTA Multi-view Images of Coronary Arteries
Antonio Tejero-de-Pablos, Kaikai Huang, Hiroaki Yamane, Yusuke Kurose, Yusuke Mukuta, Junichi Iho, Youji Tokunaga, Makoto Horie, Keisuke Nishizawa, Yusaku Hayashi, Yasushi Koyama, Tatsuya Harada
MICCAI (2)12
2019 Attention Transfer (ANT) Network for View-invariant Action Recognition
abstract
With wide applications in surveillance and human-robot interaction, view-invariant human action recognition is critical, however, challenging, due to the action occlusion and information loss caused by view change. Current methods mainly seek for a common feature space for different views. However, such solutions become invalid when there exist few common features, e.g. large view change. To tackle the problem, we propose an AttentioN Transfer (ANT) Network for view-invariant action recognition. Other than transferring features, ANT transfers attention from the reference view to arbitrary views, which correctly emphasize crucial body joints and their relations for view-invariant representation. In addition, the attention calculation method taking into account both recognition contribution and reliability of skeleton joints generates effective attention. Experiments showed its effectiveness for correctly locating crucial body joints in action sequences. We exhaustively evaluate our approach on the UESTC and the NTU dataset with three types of view-invariant evaluations, i.e. X-view, X-sub, and Arbitrary-view evaluation. Experiment results demonstrate its superiority in view-invariant representation and recognition.
Yanli Ji, Feixiang Xu, Yang Yang 0002, Ning Xie 0003, Heng Tao Shen, Tatsuya Harada
ACM Multimedia6
2019 How narratives move your mind: A corpus of shared-character stories for connecting emotional flow and interestingness
Yusuke Mori 0001, Hiroaki Yamane, Yoshitaka Ushiku, Tatsuya Harada
Inf. Process. Manag.4
2018 Alternating Circulant Random Features for Semigroup Kernels
Yusuke Mukuta, Yoshitaka Ushiku, Tatsuya Harada
AAAI3
2018 Hierarchical Video Generation From Orthogonal Information: Optical Flow and Texture
abstract
Learning to represent and generate videos from unlabeled data is a very challenging problem. To generate realistic videos, it is important not only to ensure that the appearance of each frame is real, but also to ensure the plausibility of a video motion and consistency of a video appearance in the time direction. The process of video generation should be divided according to these intrinsic difficulties. In this study, we focus on the motion and appearance information as two important orthogonal components of a video, and propose Flow-and-Texture-Generative Adversarial Networks (FTGAN) consisting of FlowGAN and TextureGAN. In order to avoid a huge annotation cost, we have to explore a way to learn from unlabeled data. Thus, we employ optical flow as motion information to generate videos. FlowGAN generates optical flow, which contains only the edge and motion of the videos to be begerated. On the other hand, TextureGAN specializes in giving a texture to optical flow generated by FlowGAN. This hierarchical approach brings more realistic videos with plausible motion and appearance consistency. Our experiments show that our model generates more plausible motion videos and also achieves significantly improved performance for unsupervised action classification in comparison to previous GAN works. In addition, because our model generates videos from two independent information, our model can generate new combinations of motion and attribute that are not seen in training data, such as a video in which a person is doing sit-up in a baseball ground.
Katsunori Ohnishi, Shohei Yamamoto, Yoshitaka Ushiku, Tatsuya Harada
AAAI4
2018 Viewpoint-Aware Video Summarization
abstract
This paper introduces a novel variant of video summarization, namely building a summary that depends on the particular aspect of a video the viewer focuses on. We refer to this as viewpoint. To infer what the desired viewpoint may be, we assume that several other videos are available, especially groups of videos, e.g., as folders on a person's phone or laptop. The semantic similarity between videos in a group vs. the dissimilarity between groups is used to produce viewpoint-specific summaries. For considering similarity as well as avoiding redundancy, output summary should be (A) diverse, (B) representative of videos in the same group, and (C) discriminative against videos in the different groups. To satisfy these requirements (A)-(C) simultaneously, we proposed a novel video summarization method from multiple groups of videos. Inspired by Fisher's discriminant criteria, it selects summary by optimizing the combination of three terms (a) inner-summary, (b) inner-group, and (c) between-group variances defined on the feature representation of summary, which can simply represent (A)-(C). Moreover, we developed a novel dataset to investigate how well the generated summary reflects the underlying viewpoint. Quantitative and qualitative experiments conducted on the dataset demonstrate the effectiveness of proposed method.
Atsushi Kanehira, Luc Van Gool, Yoshitaka Ushiku, Tatsuya Harada
CVPR4
2018 Neural 3D Mesh Renderer
abstract
For modeling the 3D world behind 2D images, which 3D representation is most appropriate? A polygon mesh is a promising candidate for its compactness and geometric properties. However, it is not straightforward to model a polygon mesh from 2D images using neural networks because the conversion from a mesh to an image, or rendering, involves a discrete operation called rasterization, which prevents back-propagation. Therefore, in this work, we propose an approximate gradient for rasterization that enables the integration of rendering into neural networks. Using this renderer, we perform single-image 3D mesh reconstruction with silhouette image supervision and our system outperforms the existing voxel-based approach. Additionally, we perform gradient-based 3D mesh editing operations, such as 2D-to-3D style transfer and 3D DeepDream, with 2D supervision for the first time. These applications demonstrate the potential of the integration of a mesh renderer into neural networks and the effectiveness of our proposed renderer.
Hiroharu Kato, Yoshitaka Ushiku, Tatsuya Harada
CVPR3
2018 Maximum Classifier Discrepancy for Unsupervised Domain Adaptation
abstract
In this work, we present a method for unsupervised domain adaptation. Many adversarial learning methods train domain classifier networks to distinguish the features as either a source or target and train a feature generator network to mimic the discriminator. Two problems exist with these methods. First, the domain classifier only tries to distinguish the features as a source or target and thus does not consider task-specific decision boundaries between classes. Therefore, a trained generator can generate ambiguous features near class boundaries. Second, these methods aim to completely match the feature distributions between different domains, which is difficult because of each domain's characteristics. To solve these problems, we introduce a new approach that attempts to align distributions of source and target by utilizing the task-specific decision boundaries. We propose to maximize the discrepancy between two classifiers' outputs to detect target samples that are far from the support of the source. A feature generator learns to generate target features near the support to minimize the discrepancy. Our method outperforms other methods on several datasets of image classification and semantic segmentation. The codes are available at https://github.com/mil-tokyo/MCD_DA.
Kuniaki Saito, Kohei Watanabe, Yoshitaka Ushiku, Tatsuya Harada
CVPR4
2018 Customized Image Narrative Generation via Interactive Visual Question Generation and Answering
abstract
Image description task has been invariably examined in a static manner with qualitative presumptions held to be universally applicable, regardless of the scope or target of the description. In practice, however, different viewers may pay attention to different aspects of the image, and yield different descriptions or interpretations under various contexts. Such diversity in perspectives is difficult to derive with conventional image description techniques. In this paper, we propose a customized image narrative generation task, in which the users are interactively engaged in the generation process by providing answers to the questions. We further attempt to learn the user's interest via repeating such interactive stages, and to automatically reflect the interest in descriptions for new images. Experimental results demonstrate that our model can generate a variety of descriptions from single image that cover a wider range of topics than conventional models, while being customizable to the target user of interaction.
Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada
CVPR3
2018 Between-Class Learning for Image Classification
abstract
In this paper, we propose a novel learning method for image classification called Between-Class learning (BC learning)1. We generate between-class images by mixing two images belonging to different classes with a random ratio. We then input the mixed image to the model and train the model to output the mixing ratio. BC learning has the ability to impose constraints on the shape of the feature distributions, and thus the generalization ability is improved. BC learning is originally a method developed for sounds, which can be digitally mixed. Mixing two image data does not appear to make sense; however, we argue that because convolutional neural networks have an aspect of treating input data as waveforms, what works on sounds must also work on images. First, we propose a simple mixing method using internal divisions, which surprisingly proves to significantly improve performance. Second, we propose a mixing method that treats the images as waveforms, which leads to a further improvement in performance. As a result, we achieved 19.4% and 2.26% top-1 errors on ImageNet-1K and CIFAR10, respectively.
Yuji Tokozume, Yoshitaka Ushiku, Tatsuya Harada
CVPR3
2018 Open Set Domain Adaptation by Backpropagation
Kuniaki Saito, Shohei Yamamoto, Yoshitaka Ushiku, Tatsuya Harada
ECCV (5)4
2018 Visual Question Generation for Class Acquisition of Unknown Objects
Kohei Uehara, Antonio Tejero-de-Pablos, Yoshitaka Ushiku, Tatsuya Harada
ECCV (12)4
2018 Adversarial Dropout Regularization
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, Kate Saenko
ICLR (Poster)3
2018 Learning from Between-class Examples for Deep Sound Recognition
Yuji Tokozume, Yoshitaka Ushiku, Tatsuya Harada
ICLR (Poster)3
2017 Learning environmental sounds with end-to-end convolutional neural network
abstract
Environmental sound classification (ESC) is usually conducted based on handcrafted features such as the log-mel feature. Meanwhile, end-to-end classification systems perform feature extraction jointly with classification and have achieved success particularly in image classification. In the same manner, if environmental sounds could be directly learned from the raw waveforms, we would be able to extract a new feature effective for classification that could not have been designed by humans, and this new feature could improve the classification performance. In this paper, we propose a novel end-to-end ESC system using a convolutional neural network (CNN). The classification accuracy of our system on ESC-50 is 5.1% higher than that achieved when using logmel-CNN with the static log-mel feature. Moreover, we achieve a 6.5% improvement in classification accuracy over the state-of-the-art logmel-CNN with the static and delta log-mel feature, simply by combining our system and logmel-CNN.
Yuji Tokozume, Tatsuya Harada
ICASSP2
2017 Spatio-Temporal Person Retrieval via Natural Language Queries
abstract
In this paper, we address the problem of spatio-temporal person retrieval from videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query. For this problem, we introduce a novel dataset consisting of videos containing people annotated with bounding boxes for each second and with five natural language descriptions. To retrieve the tube of the person described by a given natural language query, we design a model that combines methods for spatio-temporal human detection and multimodal retrieval. We conduct comprehensive experiments to compare a variety of tube and text representations and multimodal retrieval methods, and present a strong baseline in this task as well as demonstrate the efficacy of our tube representation and multimodal feature embedding technique. Finally, we demonstrate the versatility of our model by applying it to two other important tasks.
Masataka Yamaguchi, Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada
ICCV4
2017 DualNet: Domain-invariant network for visual question answering
abstract
Visual question answering (VQA) tasks use two types of images: abstract (illustrations) and real. Domain-specific differences exist between the two types of images with respect to “objectness,” “texture,” and “color.” Therefore, achieving similar performance by applying methods developed for real images to abstract images, and vice versa, is difficult. This is a critical problem in VQA, because image features are crucial clues for correctly answering the questions about the images. However, an effective, domain-invariant method can provide insight into the high-level reasoning required for VQA. We thus propose a method called DualNet that demonstrates performance that is invariant to the differences in real and abstract scene domains. Experimental results show that DualNet outperforms state-of-the-art methods, especially for the abstract images category.
Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada
ICME4
2017 Asymmetric Tri-training for Unsupervised Domain Adaptation
abstract
It is important to apply models trained on a large number of labeled samples to different domains because collecting many labeled samples in various domains is expensive. To learn discriminative representations for the target domain, we assume that artificially labeling the target samples can result in a good representation. Tri-training leverages three classifiers equally to provide pseudo-labels to unlabeled samples; however, the method does not assume labeling samples generated from a different domain. In this paper, we propose the use of an asymmetric tri-training method for unsupervised domain adaptation, where we assign pseudo-labels to unlabeled samples and train the neural networks as if they are true labels. In our work, we use three networks asymmetrically, and by asymmetric, we mean that two networks are used to label unlabeled target samples, and one network is trained by the pseudo-labeled samples to obtain target-discriminative representations. Our proposed method was shown to achieve a state-of-the-art performance on the benchmark digit recognition datasets for domain adaptation.
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada
ICML3
2017 MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes
abstract
This work addresses the semantic segmentation of images of street scenes for autonomous vehicles based on a new RGB-Thermal dataset, which is also introduced in this paper. An increasing interest in self-driving vehicles has brought the adaptation of semantic segmentation to self-driving systems. However, recent research relating to semantic segmentation is mainly based on RGB images acquired during times of poor visibility at night and under adverse weather conditions. Furthermore, most of these methods only focused on improving performance while ignoring time consumption. The aforementioned problems prompted us to propose a new convolutional neural network architecture for multi-spectral image segmentation that enables the segmentation accuracy to be retained during real-time operation. We benchmarked our method by creating an RGB-Thermal dataset in which thermal and RGB images are combined. We showed that the segmentation accuracy was significantly increased by adding thermal infrared information.
Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, Tatsuya Harada
IROS5
2017 WebDNN: Fastest DNN Execution Framework on Web Browser
abstract
Recently, deep neural network (DNN) is drawing a lot of attention because of its applications. However, it requires a lot of computational resources and tremendous processes in order to setup an execution environment based on hardware acceleration such as GPGPU. Therefore, providing DNN applications to end-users is very hard. To solve this problem, we have developed an installation-free web browser-based DNN execution framework, WebDNN. WebDNN optimizes the trained DNN model to compress model data and accelerate the execution. It executes the DNN model with novel JavaScript API to achieve zero-overhead execution. Empirical evaluations show that it achieves more than two-hundred times the unusual acceleration. WebDNN is an open source framework and you can download it from https://github.com/mil-tokyo/webdnn.
Masatoshi Hidaka, Yuichiro Kikura, Yoshitaka Ushiku, Tatsuya Harada
ACM Multimedia4
2017 Online growing neural gas for anomaly detection in changing surveillance scenes
Qianru Sun, Hong Liu 0008, Tatsuya Harada
Pattern Recognit.3
2016 Image Captioning with Sentiment Terms via Weakly-Supervised Sentiment Dataset
Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada
BMVC3
2016 Multi-label Ranking from Positive and Unlabeled Data
abstract
In this paper, we specifically examine the training of a multi-label classifier from data with incompletely assigned labels. This problem is fundamentally important in many multi-label applications because it is almost impossible for human annotators to assign a complete set of labels, although their judgments are reliable. In other words, a multilabel dataset usually has properties by which (1) assigned labels are definitely positive and (2) some labels are absent but are still considered positive. Such a setting has been studied as a positive and unlabeled (PU) classification problem in a binary setting. We treat incomplete label assignment problems as a multi-label PU ranking, which is an extension of classical binary PU problems to the wellstudied rank-based multi-label classification. We derive the conditions that should be satisfied to cancel the negative effects of label incompleteness. Our experimentally obtained results demonstrate the effectiveness of these conditions.
Atsushi Kanehira, Tatsuya Harada
CVPR2
2016 Kernel Approximation via Empirical Orthogonal Decomposition for Unsupervised Feature Learning
abstract
Kernel approximation methods are important tools for various machine learning problems. There are two major methods used to approximate the kernel function: the Nyström method and the random features method. However, the Nyström method requires relatively high-complexity post-processing to calculate a solution and the random features method does not provide sufficient generalization performance. In this paper, we propose a method that has good generalization performance without high-complexity postprocessing via empirical orthogonal decomposition using the probability distribution estimated from training data. We provide a bound for the approximation error of the proposed method. Our experiments show that the proposed method is better than the random features method and comparable with the Nyström method in terms of the approximation error and classification accuracy. We also show that hierarchical feature extraction using our kernel approximation demonstrates better performance than the existing methods.
Yusuke Mukuta, Tatsuya Harada
CVPR2
2016 Recognizing Activities of Daily Living with a Wrist-Mounted Camera
abstract
We present a novel dataset and a novel algorithm for recognizing activities of daily living (ADL) from a first-person wearable camera. Handled objects are crucially important for egocentric ADL recognition. For specific examination of objects related to users' actions separately from other objects in an environment, many previous works have addressed the detection of handled objects in images captured from head-mounted and chest-mounted cameras. Nevertheless, detecting handled objects is not always easy because they tend to appear small in images. They can be occluded by a user's body. As described herein, we mount a camera on a user's wrist. A wrist-mounted camera can capture handled objects at a large scale, and thus it enables us to skip the object detection process. To compare a wrist-mounted camera and a head-mounted camera, we also developed a novel and publicly available dataset1that includes videos and annotations of daily activities captured simultaneously by both cameras. Additionally, we propose a discriminative video representation that retains spatial and temporal information after encoding the frame descriptors extracted by convolutional neural networks (CNN).
Katsunori Ohnishi, Atsushi Kanehira, Asako Kanezaki, Tatsuya Harada
CVPR4
2016 Beyond caption to narrative: Video captioning with multiple sentences
abstract
Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image captioning process and does not fully take advantage of dynamic contents present in videos. We attempt to generate video captions that convey richer contents by temporally segmenting the video with action localization, generating multiple captions from multiple frames, and connecting them with natural language processing techniques, in order to generate a story-like caption. We show that our proposed method can generate captions that are richer in contents and can compete with state-of-the-art method wRecent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image captioning process and does not fully take advantage of dynamic contents present in videos. We attempt to generate video captions that convey richer contents by temporally segmenting the video with action localization, generating multiple captions from multiple frames, and connecting them with natural language processing techniques, in order to generate a story-like caption. We show that our proposed method can generate captions that are richer in contents and can compete with state-of-the-art method without explicitly using video-level features as input.
Andrew Shin, Katsunori Ohnishi, Tatsuya Harada
ICIP3
2016 IBC127: Video dataset for fine-grained bird classification
abstract
Beyond general object recognition whereby general categories such as dogs and cats are estimated from images, Fine-grained visual categorization (FGVC) is a new trend that goes beyond general object recognition - where general categories such as cats and dogs are estimated from images - to classify fine-grained categories of objects (or animals) such as poodles or bulldogs. It is difficult to distinguish between categories with similar appearance, e.g., sparrows and hummingbirds, using image features alone. Consequently, motion features extracted from videos are effective for classifying similar animals. In this paper, we demonstrate the effectiveness of motion features for FGVC of animals. We use our novel dataset that consists of videos of fine-grained categories of birds collected from the internet. Our dataset is publicly available for academics.
Tomoaki Saito, Asako Kanezaki, Tatsuya Harada
ICME3
2016 True-negative label selection for large-scale multi-label learning
abstract
In this paper, we focus on training a classifier from large-scale data with incompletely assigned labels. In other words, we treat samples with following properties: 1. assigned labels are definitely positive, 2. absent labels are not necessarily negative, and 3. samples are allowed to take more than one label. These properties are frequently found in various kinds of computer vision tasks, including image and video classification and retrieval. Many online algorithms for multi-label task employ label sampling, which selects a label pair that reduces the largest penalty to update the model, thereby avoiding waste of computation. In the setting above, however, there are “false-negative” labels, which are originally positive labels but regarded as negative. Since it is high likely for label sampling to select these labels as negative labels in the sampled pair, it may severely degrade classification performance. In order to solve this problem while preserving convergence property of the online algorithms, we propose a novel label sampling approach, which aims to fetch “true-negative” labels via false-negativeness measure based on independently trained uni-class classifiers. Experimental results show the effectiveness of our approach.
Atsushi Kanehira, Andrew Shin, Tatsuya Harada
ICPR3
2016 Scene Image Synthesis from Natural Sentences Using Hierarchical Syntactic Analysis
abstract
Synthesizing a new image from verbal information is a challenging task that has a number of applications. Most research on the issue has attempted to address this question by providing external clues, such as sketches. However, no study has been able to successfully handle various sentences for this purpose without any other information. We propose a system to synthesize scene images solely from sentences. Input sentences are expected to be complete sentences with visualizable objects. Our priorities are the analysis of sentences and the correlation of information between input sentences and visible image patches. A hierarchical syntactic parser is developed for sentence analysis, and a combination of lexical knowledge and corpus statistics is designed for word correlation. The entire system was applied to both a clip-art dataset and an actual image dataset. This application highlighted the capability of the proposed system to generate novel images as well as its ability to succinctly convey ideas.
Tetsuaki Mano, Hiroaki Yamane, Tatsuya Harada
ACM Multimedia3
2016 Improved Dense Trajectory with Cross Streams
abstract
Improved dense trajectories (iDT) have shown great performance in action recognition, and their combination with the two-stream approach has achieved state-of-the-art performance. It is, however, difficult for iDT to completely remove background trajectories from video with camera shaking. Trajectories in less discriminative regions should be given modest weights in order to create more discriminative local descriptors for action recognition. In addition, the two-stream approach, which learns appearance and motion information separately, cannot focus on motion in important regions when extracting features from spatial convolutional layers of the appearance network, and vice versa. In order to address the above mentioned problems, we propose a new local descriptor that pools a new convolutional layer obtained from crossing two networks along iDT. This new descriptor is calculated by applying discriminative weights learned from one network to a convolutional layer of the other network. Our method has achieved state-of-the-art performance on ordinal action recognition datasets, 92.3% on UCF101, and 66.2% on HMDB51.
Katsunori Ohnishi, Masatoshi Hidaka, Tatsuya Harada
ACM Multimedia3
2016 Video Generation Using 3D Convolutional Neural Network
abstract
Recently, content generation using neural network has been widely studied. Motivated by this recent progress, we studied the generation of videos using only a label as input. In our method, we iteratively minimize two objective functions at the same time : an objective function to evaluate how close the video is to the target class and another to evaluate how natural-appearing the video is. Our proposed method uses the cross-entropy error between the target label and the output of 3D convolutional neural network (C3D) as the objective function for evaluating how close the video is to the target class and uses the Euclidean distance between the input video and the video decoded from our temporal convolutional auto-encoder ("tempCAE") as the objective function for evaluating how natural-appearing the video is. We conducted an experiment evaluating the generated videos using a crowdsourcing service and confirmed the utility of our method.
Shohei Yamamoto, Tatsuya Harada
ACM Multimedia2
2015 Common Subspace for Model and Similarity: Phrase Learning for Caption Generation from Images
abstract
Generating captions to describe images is a fundamental problem that combines computer vision and natural language processing. Recent works focus on descriptive phrases, such as "a white dog" to explain the visual composites of an input image. The phrases can not only express objects, attributes, events, and their relations but can also reduce visual complexity. A caption for an input image can be generated by connecting estimated phrases using a grammar model. However, because phrases are combinations of various words, the number of phrases is much larger than the number of single words. Consequently, the accuracy of phrase estimation suffers from too few training samples per phrase. In this paper, we propose a novel phrase-learning method: Common Subspace for Model and Similarity (CoSMoS). In order to overcome the shortage of training samples, CoSMoS obtains a subspace in which (a) all feature vectors associated with the same phrase are mapped as mutually close, (b) classifiers for each phrase are learned, and (c) training samples are shared among co-occurring phrases. Experimental results demonstrate that our system is more accurate than those in earlier work and that the accuracy increases when the dataset from the web increases.
Yoshitaka Ushiku, Masataka Yamaguchi, Yusuke Mukuta, Tatsuya Harada
ICCV4
2015 3D Selective Search for obtaining object candidates
abstract
We propose a new method for obtaining object candidates in 3D space. Our method requires no learning, has no limitation of object properties such as compactness or symmetry, and therefore produces object candidates using a completely general approach. This method is a simple combination of Selective Search, which is a non-learning-based objectness detector working in 2D images, and a supervoxel segmentation method, which works with 3D point clouds. We made a small but non-trivial modification to supervoxel segmentation; it brings better “seeding” for supervoxels, which produces more proper object candidates as a result. Our experiments using a couple of publicly available RGB-D datasets demonstrated that our method outperformed state-of-the-art methods of generating object proposals in 2D images.
Asako Kanezaki, Tatsuya Harada
IROS2
2015 Probabilistic Semi-Canonical Correlation Analysis
abstract
anonical Correlation Analysis (CCA) requires paired multimodal data to ascertain the relation between two variables. However, it is generally difficult to collect a sufficient amount of paired data of two variables as training samples. This fact leads individual samples of unpaired variables to be additional resources for learning CCA, which are not only able to increase the number of training samples; they are also effective to remove the learning bias caused by the variables' missing patterns. As described in this paper, we propose a novel model of probabilistic CCA by considering the mechanism of data missing. Our method enables widespread applications such as semi-supervised learning via partially labeled training samples and analysis of sensory data which are lacking under certain circumstances. We demonstrate the superior performance of parameter estimation as well as an application of image annotation, compared with existing methods.
Chie Kamada, Asako Kanezaki, Tatsuya Harada
ACM Multimedia3
2014 Learning Similarities for Rigid and Non-rigid Object Detection
abstract
In this paper, we propose an optimization method for estimating the parameters that typically appear in graph-theoretical formulations of the matching problem for object detection. Although several methods have been proposed to optimize parameters for graph matching in a way to promote correct correspondences and to restrict wrong ones, our approach is novel in the sense that it aims at improving performance in the more general task of object detection. In our formulation, similarity functions are adjusted so as to increase the overall similarity among a reference model and the observed target, and at the same time reduce the similarity among reference and "non-target" objects. We evaluate the proposed method in two challenging scenarios, namely object detection using data captured with a Kinect sensor in a real environment, and intrinsic metric learning for deformable shapes, demonstrating substantial improvements in both settings.
Asako Kanezaki, Emanuele Rodolà, Daniel Cremers, Tatsuya Harada
3DV4
2014 Image Reconstruction from Bag-of-Visual-Words
abstract
The objective of this study is to reconstruct images from Bag-of-Visual-Words (BoVW), which is the de facto standard feature for image retrieval and recognition. BoVW is defined here as a histogram of quantized descriptors extracted densely on a regular grid at a single scale. Despite its wide use, no report describes reconstruction of the original image of a BoVW. This task is challenging for two reasons: 1) BoVW includes quantization errors when local descriptors are assigned to visual words. 2) BoVW lacks spatial information of local descriptors when we count the occurrence of visual words. To tackle this difficult task, we use a large-scale image database to estimate the spatial arrangement of local descriptors. Then this task creates a jigsaw puzzle problem with adjacency and global location costs of visual words. Solving this optimization problem is also challenging because it is known as an NP-Hard problem. We propose a heuristic but efficient method to optimize it. To underscore the effectiveness of our method, we apply it to BoVWs extracted from about 100 different categories and demonstrate that it can reconstruct the original images, although the image features lack spatial information and include quantization errors.
Hiroharu Kato, Tatsuya Harada
CVPR2
2014 Three Guidelines of Online Learning for Large-Scale Visual Recognition
abstract
In this paper, we would like to evaluate online learning algorithms for large-scale visual recognition using state-of-the-art features which are preselected and held fixed. Today, combinations of high-dimensional features and linear classifiers are widely used for large-scale visual recognition. Numerous so-called mid-level features have been developed and mutually compared on an experimental basis. Although various learning methods for linear classification have also been proposed in the machine learning and natural language processing literature, they have rarely been evaluated for visual recognition. Therefore, we give guidelines via investigations of state-of-the-art online learning methods of linear classifiers. Many methods have been evaluated using toy data and natural language processing problems such as document classification. Consequently, we gave those methods a unified interpretation from the viewpoint of visual recognition. Results of controlled comparisons indicate three guidelines that might change the pipeline for visual recognition.
Yoshitaka Ushiku, Masatoshi Hidaka, Tatsuya Harada
CVPR3
2014 Mirror reflection invariant HOG descriptors for object detection
abstract
Histogram of Oriented Gradients (HOG) [1] descriptors have been widely used for object detection. An important limitation is that these descriptors tend to vary considerably when objects are horizontally flipped, as is often the case. We propose novel MI-HOG descriptors that are obtained by transforming HOG descriptors to be invariant to mirror reflection. In their extraction process, we consider not only the transform of independent elements but also the combination of those in different location and in orientation, which yields better performance. We showed a greater than 10 % increase in average precision compared to HOG descriptors.
Asako Kanezaki, Yusuke Mukuta, Tatsuya Harada
ICIP3
2014 Probabilistic Partial Canonical Correlation Analysis
abstract
Partial canonical correlation analysis (partial CCA) is a statistical method that estimates a pair of linear projections onto a low dimensional space, where the correlation between two multidimensional variables is maximized after eliminating the influence of a third variable. Partial CCA is known to be closely related to a causality measure between two time series. However, partial CCA requires the inverses of covariance matrices, so the calculation is not stable. This is particularly the case for high-dimensional data or small sample sizes. Additionally, we cannot estimate the optimal dimension of the subspace in the model. In this paper, we have addressed these problems by proposing a probabilistic interpretation of partial CCA and deriving a Bayesian estimation method based on the probabilistic model. Our numerical experiments demonstrated that our methods can stably estimate the model parameters, even in high dimensions or when there are a small number of samples.
Yusuke Mukuta, Tatsuya Harada
ICML2
2014 Hard negative classes for multiple object detection
abstract
We propose an efficient method to train multiple object detectors simultaneously using a large scale image dataset. The one-vs-all approach that optimizes the boundary between positive samples from a target class and negative samples from the others has been the most standard approach for object detection. However, because this approach trains each object detector independently, the scores are not balanced between object classes. The proposed method combines ideas derived from both detection and classification in order to balance the scores across all object classes. We optimized the boundary between target classes and their “hard negative” samples, just as in detection, while simultaneously balancing the detector scores across object classes, as done in multi-class classification. We evaluated the performances on multi-class object detection using a subset of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2011 dataset and showed our method outperformed a de facto standard method.
Asako Kanezaki, Sho Inaba, Yoshitaka Ushiku, Yuya Yamashita, Hiroshi Muraoka, Yasuo Kuniyoshi, Tatsuya Harada
ICRA7
2014 Automatic Image Synthesis from Keywords Using Scene Context
abstract
Text is one of the simplest way to express one's idea, and an image is one of the most impactive way to do so. Therefore, if a system can synthesize an image from text without direct user manipulation, novel image synthesis applications will be opened to users without artistic skills. In such a system, which objects to synthesize will be declared in texts. However, information about positional relations and scale of objects is not much provided and must be estimated using common sense. As described in this paper, we develop a system that can automatically synthesize objects to an image, given the background image and class name of the target synthesizing object. With the inputs as the background image and keywords, images for synthesizing objects are searched automatically. Although some previously developed systems that can synthesize an image from sketches and paintings, this is the first system that can estimate the position, scale, and appearance of objects and automatically synthesize them to images without direct user input. We propose a scene context, which indicates the position, scale, and appearance of synthesizing objects. The contribution of this paper is twofold: (1) the scene context extraction method for automatic image synthesis and (2) application of automatic image synthesis using the scene context.
Sho Inaba, Asako Kanezaki, Tatsuya Harada
ACM Multimedia3
2014 Clothing Retrieval Based on Local Similarity with Multiple Images
abstract
Recently, the online shopping market has been expanded, which has advanced studies of clothing retrieval via image search. For this study, we develop a novel clothing retrieval system considering local similarity, where users can retrieve their desired clothes which are globally similar to an image and partially similar to another image. We propose a method of coding global features by merging local descriptors extracted from multiple images. Furthermore, we design a system that re-evaluates output of similar image search by the similarity of local regions. We demonstrated that our method increased the probability of users finding their desired clothes from 39.7%-55.1%, compared to a standard similar image search system with global features of a single image. Statistical significance is proven using t-tests.
Masaru Mizuochi, Asako Kanezaki, Tatsuya Harada
ACM Multimedia3
2013 Efficient Shape Matching using Vector Extrapolation
abstract
We propose the adoption of a vector extrapolation technique to accelerate convergence of correspondence problems under the quadratic assignment formulation for attributed graph matching (QAP). In order to capture a broad range of matching scenarios, we provide a class of relaxations of the QAP under elastic net constraints. This allows us to regulate the sparsity/complexity trade-off which is inherent to most instances of the matching problem, thus enabling us to study the application of the acceleration method over a family of problems of varying difficulty. The validity of the approach is assessed by considering three different matching scenarios; namely, rigid and non-rigid three-dimensional shape matching, and image matching for Structure from Motion. As demonstrated on both real and synthetic data, our approach leads to an increase in performance of up to one order of magnitude when compared to the standard methods. 1
Emanuele Rodolà, Tatsuya Harada, Yasuo Kuniyoshi, Daniel Cremers
BMVC2
2013 Elastic Net Constraints for Shape Matching
abstract
We consider a parametrized relaxation of the widely adopted quadratic assignment problem (QAP) formulation for minimum distortion correspondence between deformable shapes. In order to control the accuracy/sparsity trade-off we introduce a weighting parameter on the combination of two existing relaxations, namely spectral and game-theoretic. This leads to the introduction of the elastic net penalty function into shape matching problems. In combination with an efficient algorithm to project onto the elastic net ball, we obtain an approach for deformable shape matching with controllable sparsity. Experiments on a standard benchmark confirm the effectiveness of the approach.
Emanuele Rodolà, Andrea Torsello, Tatsuya Harada, Yasuo Kuniyoshi, Daniel Cremers
ICCV3
2013 Weakly-supervised multi-class object detection using multi-type 3D features
abstract
We propose a weakly-supervised learning method for object detection using color and depth images of a real environment attached with object labels. The proposed method applies Multiple Instance Learning to find proper instances of the objects in training images. This method is novel in the sense that it learns multiple objects simultaneously in a way to balance the scores of each training sample across all object classes. Moreover, we combine 3D features considering different properties, that is, color texture, grayscale texture, and surface curvature, to improve the performance. We show that our method surpasses a conventional method using color and depth images. Furthermore, we evaluate its performance with our new dataset consisting of color and depth images with weak labels of 100 objects and various backgrounds.
Asako Kanezaki, Yasuo Kuniyoshi, Tatsuya Harada
ACM Multimedia3
2012 Visual anomaly detection from small samples for mobile robots
abstract
We propose a novel method of visual anomaly detection for mobile robots in daily real-life settings. Visual anomaly detection using mobile robots is important for security systems or simply for gathering information. However, this task is challenging for two reasons. First, because the number of observed images sampled at the same location is small, anomaly detection systems cannot use standard statistical methods. Second, anomalies must be detected in the presence of other continuous, ambient changes in the visual scene, such as changes in lighting from morning to night. Regarding the former problem, we develop and apply an analysis-by-synthesis-based anomaly detection method for mobile robots. For the latter, we propose a novel definition of anomaly that uses observed samples at other locations to filter out ambient changes that should be ignored by the system. Experimental results demonstrate that our method can detect anomalies from small samples in the presence of ambient changes, which could not be detected by conventional methods.
Hiroharu Kato, Tatsuya Harada, Yasuo Kuniyoshi
IROS2
2012 Efficient image annotation for automatic sentence generation
abstract
Sentence generation from images is an ultimate goal of image recognition. In this paper, we attack a novel problem, the "multi-keyphrase problem", to address this goal. We hypothesize that image contents can be described with multi-keyphrases, and that a natural sentence can be generated by connecting multi-keyphrases with an experimental grammar model. Existing methods require semantic knowledge such as labels of an object, action, or scene. Using these methods, we must strive to prepare a highly organized dataset. Therefore, we propose a novel online learning method for multi-keyphrase estimation. The proposed framework, although simple and scalable, can generate sentences from images with no semantic knowledge. Moreover, the proposed method for multi-keyphrase estimation is applicable to image annotation, and it achieves state-of-the-art performance. Our experiment using only images and texts demonstrates that the proposed framework is useful for sentence generation from images.
Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi
ACM Multimedia2
2012 Graphical Gaussian Vector for Image Categorization
abstract
This paper proposes a novel image representation called a Graphical Gaussian Vector, which is a counterpart of the codebook and local feature matching approaches. In our method, we model the distribution of local features as a Gaussian Markov Random Field (GMRF) which can efficiently represent the spatial relationship among local features. We consider the parameter of GMRF as a feature vector of the image. Using concepts of information geometry, proper parameters and a metric from the GMRF can be obtained. Finally we define a new image feature by embedding the metric into the parameters, which can be directly applied to scalable linear classifiers. Our method obtains superior performance over the state-of-the-art methods in the standard object recognition datasets and comparable performance in the scene dataset. As the proposed method simply calculates the local auto-correlations of local features, it is able to achieve both high classification accuracy and high efficiency.
Tatsuya Harada, Yasuo Kuniyoshi
NIPS1
2012 Dialog System Using Real-Time Crowdsourcing and Twitter Large-Scale Corpus
Fumihiro Bessho, Tatsuya Harada, Yasuo Kuniyoshi
SIGDIAL Conference2
2012 Causal Flow
abstract
Optical flow is a widely used technique for extracting flow information from video images. While it is useful for estimating temporary movement in video images, it only captures one aspect of extracting dominant flow information from a sequence of video images. In this paper, we propose a novel flow extraction approach called causal flow, which can estimate the dominant causal relationships among nearby pixels. We assume flows in video images as pixel-to-pixel information transfer, whereas the optical flow measures the relative motion of pixels. Causal flow is based on the Granger causality test, which measures causal influence based on prediction via vector autoregression, and is widely used in economics and brain science. The experimental results demonstrate that causal flow can extract dominant flow information which cannot be obtained by current methods.
Yuya Yamashita, Tatsuya Harada, Yasuo Kuniyoshi
IEEE Trans. Multim.2
2011 Discriminative spatial pyramid
abstract
Spatial Pyramid Representation (SPR) is a widely used method for embedding both global and local spatial information into a feature, and it shows good performance in terms of generic image recognition. In SPR, the image is divided into a sequence of increasingly finer grids on each pyramid level. Features are extracted from all of the grid cells and are concatenated to form one huge feature vector. As a result, expensive computational costs are required for both learning and testing. Moreover, because the strategy for partitioning the image at each pyramid level is designed by hand, there is weak theoretical evidence of the appropriate partitioning strategy for good categorization. In this paper, we propose discriminative SPR, which is a new representation that forms the image feature as a weighted sum of semi-local features over all pyramid levels. The weights are automatically selected to maximize a discriminative power. The resulting feature is compact and preserves high discriminative power, even in low dimension. Furthermore, the discriminative SPR can suggest the distinctive cells and the pyramid levels simultaneously by observing the optimal weights generated from the fine grid cells.
Tatsuya Harada, Yoshitaka Ushiku, Yuya Yamashita, Yasuo Kuniyoshi
CVPR1
2011 Causal flow
abstract
Optical flow is a widely used technique for extracting flow information from video images. While it is useful for estimating temporary movement in video images, it only captures one aspect of extracting dominant flow information from a sequence of video images. In this paper, we propose a novel flow extraction approach called causal flow, which can estimate the dominant causal relationships among nearby pixels. We assume flows in video images as pixel-to-pixel information transfer, whereas the optical flow measures the relative motion of pixels. Causal flow is based on Granger causality test, which measures causal influence based on prediction via vector autoregression, and is widely used in economics and brain science. The experimental results demonstrate that causal flow can extract dominant flow information which cannot be obtained by current methods.
Yuya Yamashita, Tatsuya Harada, Yasuo Kuniyoshi
ICME2
2011 Fast object detection for robots in a cluttered indoor environment using integral 3D feature table
abstract
Realizing automatic object search by robots in an indoor environment is one of the most important and challenging topics in mobile robot research. If the target object does not exist in a nearby area, the obvious strategy is to go to the area in which it was last observed. We have developed a robot system that collects 3D-scene data in an indoor environment during automatic routine crawling, and also detects objects quickly through a global search of the collected 3D-scene data. The 3D-scene data can be obtained automatically by transforming color images and range images into a set of color voxel data using self-location information. To detect an object, the system moves the bounding box of the target object by a certain step in the color voxel data, extracts 3D features in each box region, and computes the similarity between these features and the target object's features, using an appropriate feature projection learned beforehand. Taking advantage of the additive property of our 3D features, both feature extraction and similarity calculation are considerably accelerated. In the object learning process, the system obtains the feature-projection matrix by weighting unique features of the target object rather than its common features, resulting in reducing object detection errors.
Asako Kanezaki, Tatsuya Harada, Yasuo Kuniyoshi
ICRA3
2011 Visual anomaly detection under temporal and spatial non-uniformity for news finding robot
abstract
In this paper, we propose a news-gathering mobile robot system, and the novel visual anomaly detection method as the core function of news detection in the real world. Visual anomaly detection is important and widely applicable not only to the news-gathering robot but also to the security systems. However, visual anomaly detection from the mobile robot is highly challenging, because the appearances of images captured by the moving robot are dynamically changing. In consequence, the number of observed images at the same location becomes small, and the sampling interval of those images is not constant. To tackle this problem, we developed a new method to incorporate many samples observed at different locations as previous knowledge, which implicitly represent semantically similar to the intended location. Also, we developed a new statistical model, which explicitly considers sampling interval of input images, whereas conventional methods ignore correlation among samples. Experimental results demonstrate that our method outperforms conventional methods, and our mobile robot system including the proposed method finds, investigates, and publishes news of a local community of the real world.
Fumihiro Bessho, Tatsuya Harada, Yasuo Kuniyoshi
IROS3
2011 Efficient multi-modal retrieval in conceptual space
abstract
In this paper, we propose a new, efficient retrieval system for large-scale multi-modal data including video tracks. With large-scale multi-modal data, the huge data size and various contents cause degradation of efficiency and precision of retrieval results. Recent research on image annotation and retrieval shows that image features based on the Bag-of-Visual Words approach with local descriptors such as SIFT perform surprisingly well with large-scale image datasets. Those powerful descriptors tend to be high-dimensional, imposing a high computational cost for approximate nearest neighbor searching in raw feature space. Our video retrieval method is focused on the correlation between image, sound, and location information recorded simultaneously, and to learn conceptual space describing the contents of the data to realize efficient searching. Experiments show good performance of our retrieval system with low memory usage and temporal complexity.
Jun Imura, Teppei Fujisawa, Tatsuya Harada, Yasuo Kuniyoshi
ACM Multimedia3
2011 Understanding images with natural sentences
abstract
We propose a novel system which generates sentential captions for general images. For people to use numerous images effectively on the web, technologies must be able to explain image contents and must be capable of searching for data that users need. Moreover, images must be described with natural sentences based not only on the names of objects contained in an image but also on their mutual relations. The proposed system uses general images and captions available on the web as training data to generate captions for new images. Furthermore, because the learning cost is independent from the amount of data, the system has scalability, which makes it useful with large-scale data.
Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi
ACM Multimedia2
2011 Automatic sentence generation from images
abstract
For the overwhelming amounts of multimedia used on the Web, methods of search and understanding with sentences are necessary. Representing the contents not only using labels but also using sentences including labels' relations enables users to search with a story and to understand multimedia deeply. However, few existing works describe such sentences because obtaining objects' relations and grammar is difficult. We specifically examine captions of images that are similar to an input image. They are expected to explain the input image to some degree. Therefore, we propose a novel approach to generate a sentential caption for the input image by summarizing those captions. Our experiment using a dataset consisting of images and text demonstrates that the proposed method can generate sentential captions.
Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi
ACM Multimedia2
2010 Evaluation of dimensionality reduction methods for image auto-annotation
abstract
Image auto-annotation is a challenging task in computer vision. The goal of this task is to predict multiple words for generic images automatically. Recent state-of-theart methods are based on a non-parametric approach that uses several visual features to calculate distances between image samples. While this approach is successful from the viewpoint of annotation accuracy, the computational costs, in terms of both complexity and memory use, tend to be high, since non-parametric methods require many training instances to be stored in memory to compute distances from a query. In this paper, we investigate several linear dimensionality reduction methods for efficient image annotation. Using the additional information provided by multiple labels, we can obtain a small representation preserving (and hopefully improving) the semantic distance of a visual feature. Linear methods are computationally reasonable and are suitable for practical large-scale systems, although only limited comparison of such methods is available in this research field. Extensive experiments and analyses on various datasets and visual features show how these simple methods can be applied effectively to image annotation.
Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
BMVC2
2010 Global Gaussian approach for scene categorization using information geometry
abstract
Local features provide powerful cues for generic image recognition. An image is represented by a “bag” of local features, which form a probabilistic distribution in the feature space. The problem is how to exploit the distributions efficiently. One of the most successful approaches is the bag-of-keypoints scheme, which can be interpreted as sparse sampling of high-level statistics, in the sense that it describes a complex structure of a local feature distribution using a relatively small number of parameters. In this paper, we propose the opposite approach, dense sampling of low-level statistics. A distribution is represented by a Gaussian in the entire feature space. We define some similarity measures of the distributions based on an information geometry framework and show how this conceptually simple approach can provide a satisfactory performance, comparable to the bag-of-keypoints for scene classification tasks. Furthermore, because our method and bag-of-keypoints illustrate different statistical points, we can further improve classification performance by using both of them in kernels.
Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
CVPR2
2010 Improving Local Descriptors by Embedding Global and Local Spatial Information
Tatsuya Harada, Hideki Nakayama, Yasuo Kuniyoshi
ECCV (4)1
2010 Improving image similarity measures for image browsing and retrieval through latent space learning between images and long texts
abstract
The amount of multimedia data on personal devices and the Web is increasing daily. Image browsing and retrieval systems in a low-dimensional space have been widely studied to manage and view large numbers of images. It is essential for such systems to exploit an efficient similarity measure of the images when searching for them. Existing methods use the distance in a low-level image feature space as the similarity measure, and therefore, images with different content may be treated as similar images. In this paper, we propose a novel method to improve the similarity measures for images by considering the text surrounding the images. If there is text describing the images, similarities can be measured more effectively by taking into account the text streams. The proposed method improves the image similarity measures based on the latent semantics obtained from the combination of image and text. It should be noted that the text does not need to be clear tags; indeed, any generic Web text is applicable. Moreover, our method can effectively improve the similarities even if only a small portion of the images include textual descriptions. Additionally, the proposed method is scalable as it has linear computational complexity based on the number of images. In the experiments, we compare our method with previous methods using an original dataset in which a portion of the images are annotated by long text. We show that the proposed method can retrieve semantically similar images more precisely than existing methods.
Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi
ICIP2
2010 High-speed 3D object recognition using additive features in a linear subspace
abstract
In this paper we propose a method of high-speed 3D object recognition using linear subspace method and our 3D features. This method can be applied to partial models with any size in any posture. Although it is becoming easy to obtain textured 3D models by a 3D scanner, there are few methods for 3D object recognition which take into account both shape and textures of objects. Moreover, it is difficult to achieve high-speed processing of large 3D data. Our 3D features consider the co-occurrence of shape and colors of an object's surface. The additive property of these features makes it possible to calculate the similarity between a query part and the subspace of each object in a database without division, and therefore the time for recognition is quite short. In the experiments, we compare our method with conventional methods using Spin-Images and Textured Spin-Images. We show that our method is appropriate for 3D object recognition.
Asako Kanezaki, Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
ICRA3
2010 Partial matching of real textured 3D objects using color cubic higher-order local auto-correlation features
Asako Kanezaki, Tatsuya Harada, Yasuo Kuniyoshi
Vis. Comput.2
2009 Image annotation and retrieval based on efficient learning of contextual latent space
abstract
Image annotation and retrieval are extremely difficult because of the generic nature of the target images. Generic images contain various miscellaneous objects and scenes. Therefore, desirable annotation results are subjective and underspecified. To overcome this problem, it is important to assume "weak labeling" framework, where images are weakly related to multiple words without region information. In this paper, we propose a high speed and high accuracy image annotation and retrieval method based on efficient learning of the contextual latent space. A distance between samples can be defined in the intrinsic feature space for annotation using latent space learning between images and labels. The proposed method is shown to be faster and more accurate than previously published methods.
Tatsuya Harada, Hideki Nakayama, Yasuo Kuniyoshi
ICME1
2009 Wearable motion capture suit with full-body tactile sensors
abstract
This paper presents a system for capturing human movement and tactile data and methods for analyzing this data. We cannot fully capture the essence of motion without tactile information, and sometimes the lack of such information causes critical problems. To achieve a better understanding of motion behavior, we developed a wearable motion capture suit with full-body tactile sensors. We also developed a motion sensor which can estimate its orientation with its inner CPU. We also built a tactile sensor module which can fit many kinds of body shapes. With this system, we can measure a user's movement and tactile information simultaneously. By integrating tactile data with motion data, we can achieve many kinds of meaningful insights. We demonstrate the effectiveness of this system with experiments. We captured two motions: stretching after sitting on a chair and laying down on a bed. By recognizing the contact point from the tactile data and fitting it into the environment, we were able to estimate the motion trajectories.
Yuki Fujimori, Yoshiyuki Ohmura, Tatsuya Harada, Yasuo Kuniyoshi
ICRA3
2009 Causality quantification and its applications: structuring and modeling of multivariate time series
abstract
Time series prediction is an important issue in a wide range of areas. There are various real world processes whose states vary continuously, and those processes may have influences on each other. If the past information of one process X improves the predictability of another process Y, X is said to have a causal influence on Y. In order to make good predictions, it is necessary to identify the appropriate causal relationships. In addition, the processes to be modeled may include symbolic data as well as numerical data. Therefore, it is important to deal with symbolic and numerical time series seamlessly when attempting to detect causality.
Takashi Shibuya 0001, Tatsuya Harada, Yasuo Kuniyoshi
KDD2
2008 Smart extraction of desired object from color-distance image with user's tiny scribble
abstract
Image segmentation is an important problem because it is required for many different applications. In particular, visual extraction of an object that is the target of attention or manipulation, is an increasingly important issue in robot vision. In real-world applications, a robot needs to extract an object designated by a human in a complicated environment. There is a large literature on the problem of image segmentation, but most previous methods have a limited ability to extract desired objects from a cluttered scene. Moreover, from the perspective of human-robot interfaces, it is desirable to make it as easy as possible for the user to indicate an object. In this paper, we propose a segmentation method, CD-matting, which can correctly extract a target object in complicated real-world visual situations. This method exploits color and distance information in an integrated way. The system requires only a simple input to designate the target object. We verify the proposed system by real-world experiments. The results show the effectiveness of our method in complicated situations.
Naoki Shibuya, Yasuyuki Shimohata, Tatsuya Harada, Yasuo Kuniyoshi
IROS3
2007 Development of Wireless Networked Tiny Orientation Device for Wearable Motion Capture and Measurement of Walking Around, Walking Up and Down, and Jumping Tasks
abstract
In this paper, we developed a tiny orientation device equipped with a wireless network function for a wearable motion capture. The wearable motion capture is defined that it not only measures the posture of the human body but also collects environmental information and the human’s internal state simultaneously and easily. Because the realized device automatically configures wireless networks and is small enough to attach it anywhere, it is easy to gather any sensor information. The feature of the orientation estimation method is that models are switched according to the environment to exclude the effect of motion disturbances. In experiments, by integrating of sole sensors and the orientation sensors, walking around, walking up and down, and jumping tasks were successfully measured. Because it is difficult to measure these motions with only the inertial sensors, it showed the importance of integrating various sensors for acquiring human motion.
Tatsuya Harada, Tomoaki Gyota, Yasuo Kuniyoshi, Tomomasa Sato
IROS1
2007 Journalist robot: robot system making news articles from real world
abstract
We describe the development of a journalist robot system, which generates articles by searching for news in the real world. Our system repeats the steps: (1) autonomous exploration (2) recording of news, and (3) generation of articles. We characterize events with two values: "anomaly" and "relevance" to the user. During the exploration step, images are evaluated using these values. If an interesting event is detected, the robot approaches it to collect additional information. The system then labels the images, and generates a description from the labels. Experiments show the ability of our system to find news-like phenomena and describe images with words.
Rie Matsumoto, Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
IROS3
2006 Imitation Learning System to Assist Human Task Interactively
abstract
This paper proposes an imitation learning system to generate trajectories by which a robot supports a human with close physical assistance adapting to human movements and daily life environments. The proposed system is composed of 1) division algorithms, 2) learning algorithms and 3) assistance algorithms. 1) In division algorithms, the system measures time series of human task execution data and divides them into multiple motion segments automatically. This division is based on standard deviations of motion errors between measured trajectories and an ideal trajectory where the ideal trajectory is mean of all measured human trajectories and is expected to achieve the purpose of human task successfully. Since an important motion parameter is paid attention to by the human and has small standard deviation of errors, series of measured data are divided into segment motions at the points where the importance of parameters changes suddenly. Thus this division is guaranteed to accord with human attention. 2) In learning algorithms, the system learns trajectories with dynamic neural network (DNN). Since the DNN has convergence, generated trajectories can converge to an ideal trajectory. The importance of each parameter, in other words how much attention human pays to the parameter, is evaluated as how small the standard deviation of errors is. The DNN learns trajectories reflecting the evaluated importance of parameters to accord with human feeling. 3) In assistance algorithms, the system judges when to start assistance by the assumption of multiplied errors of motion parameters by the respective importance. In assistance algorithms, the system also connects generated trajectories of motion segments smoothly. An experiment to support human drink task was performed successfully where the proposed system judged not only when to start assistance to the task but also execute assistance when a cup was about to incline too much not to spill water
So Taoka, Tatsuya Harada, Tomomasa Sato, Taketoshi Mori
IROS2
2005 Marginalized Bags of Vectors Kernels on Switching Linear Dynamics for Online Action Recognition
abstract
In this paper, we propose a novel kernel computation algorithm between time-series human motion data for online action recognition. The proposed kernel is based on probabilistic models called switching linear dynamics (SLDs). SLD is one of the powerful tools for tracking, analyzing and classifying human complex time-series motion. The proposed kernel incorporates information about the latent variables in SLDs with simplified designing approach called marginalized kernels. The empirical evaluation using real motion data shows that a classifier using SVM with our proposed kernel has much better performance than the classifier with some conventional kernel techniques. Another experiment using walking around motion shows that a classifier with the proposed kernel can properly segment the start and the end of the target action.
Masamichi Shimosaka, Taketoshi Mori, Tatsuya Harada, Tomomasa Sato
ICRA3
2005 Construction of wireless ad hoc network for Lifelog based physical and informational support system
abstract
In this paper, we constructed wireless ad hoc network for realizing various physical and informational support systems based on Lifelog that is the records of experiences in the daily life, and realized a prototype of electric appliances operational support system which is one of useful systems. Utilization of Lifelog reduces the burden for controlling a huge amount of convoluted electric appliances by constructing the probabilistic model of user's operational behavior based on Lifelog and predicting user's successive operations with this model. In this time, Lifelog accumulates user's operational behavior for electric appliances and environmental information through the wireless network. For the basis of collecting Lifelog including operational behavior, we made a portable Bluetooth-equipped device for wireless network which is necessary to communicate information in the ubiquitous computing environment because Bluetooth has good features such as the ad hoc networking function, sufficient data throughput, high resistance to noise and low power consumption. The realized device has abundant 110 connectors which are able to connect various sensors and actuators. By attaching these devices to electric appliances, they can easily participate in the wireless network and communicate information. Therefore the system can probabilistically model user's behavior, especially operations for electric appliances using Lifelog as training data, and predicts the user's next successive operations for surrounding appliances utilizing this model and the user's present state. The system gives prediction results to the user and executes operations via the wireless network after the user's confirmation. The various experimental results prove that the operational system in ubiquitous computing environment is useful and our realized device is sufficient performance in the daily life.
Tatsuya Harada, Yusuke Kawano, Satoshi Otani, Taketoshi Mori, Tomomasa Sato
IROS1
2005 Human posture reconstruction based on posture probability density
abstract
In this paper, we propose a human posture reconstruction method from the insufficient input posture data based on human posture probability density that is constructed by a long-term human motion capture data. Since the long continuous daily human motion data has high dimensions and becomes huge size, the human posture data should be effectively compressed. The long term posture data has nonlinear distribution on the posture space, since each specific posture such as standing and sitting has different property. The posture data is allocated into some subspaces and compressed for each subspace with mixtures of probabilistic principal component analyzer (MPPCA). MPPCA is improved by replacing conventional EM algorithm with deterministic annealing EM algorithm (DAEM) to avoid initial parameter sensitivity. The posture probability density is constructed over those subspaces. The adequate human posture can be reconstructed from the insufficient data by introducing the posture probability density into the sequential Monte Carlo framework. The experimental results show that the robust human posture estimation can be realized since this method does not estimate the unique posture but estimates the proper posterior posture density with using the posture prior knowledge.
Tatsuya Harada, Tomomasa Sato, Taketoshi Mori
IROS1
2005 Online recognition and segmentation for time-series motion with HMM and conceptual relation of actions
abstract
In this paper, we propose a robust online action recognition algorithm with a segmentation scheme that detects start and end points of action occurrences. In other words, the algorithm estimates reliably what kind of actions occurring at present time. The algorithm has following characteristics: 1) The algorithm incorporates human knowledge about relation between action names in order to simplify and toughen the algorithm, thus our algorithm can label robustly multiple action names at the same time. 2) The algorithm uses time-series action probability that represents the likelihood of each action occurrence at every frame time. 3) The classification technique with hidden Markov models (HMMs) enables the algorithm to detect robustly and immediately the segmental points. The experimental results using real motion capture data show that our algorithm not only decreases effectively the latency for detecting the segmental points but also prevents the system from making unnecessary segments due to the error of time-series action probability.
Taketoshi Mori, Yu Nejigane, Masamichi Shimosaka, Yushi Segawa, Tatsuya Harada, Tomomasa Sato
IROS5
2005 Behavior prediction based on daily-life record database in distributed sensing space
abstract
This paper proposes a behavior prediction system for supporting our daily lives. The behaviors in daily-life are recorded in an environment with embedded sensors, and the prediction system learns the characteristic patterns that would be followed by the behaviors to be predicted. In this research, the authors applied a method of discovering time-series association rules, which discovers frequent combinations of events called episodes. The prediction system observes behaviors with the sensors and outputs the prediction of the future behaviors based on the rules.
Taketoshi Mori, Aritoki Takada, Hiroshi Noguchi, Tatsuya Harada, Tomomasa Sato
IROS4
2004 Portable Absolute Orientation Estimation Device with Wireless Network under Accelerated Situation
abstract
In this paper, we develop an absolute orientation estimation device equipped with a wireless network. Accelerometers and magnetometers are used to measure the gravity and the geomagnetic field respectively. Gyroscope sensors are used to measure the local angular velocity. The geomagnetic field varies according to the environment. Therefore the device can obtain the information about the magnetic field through the wireless network. The orientation estimation task can be also requested to the other computers with the wireless network. By integrating the measured gravity and geomagnetic field with the local angular velocity using Sigma-Points Kalman Filters (SPKFs), the stability and the robustness of estimating the absolute orientation are improved over either sensor alone. We also propose an estimation method which excludes the effect of motion and magnetic disturbances for the accurate estimation.
Tatsuya Harada, Hiroto Uchino, Taketoshi Mori, Tomomasa Sato
ICRA1
2004 Informative motion extractor for action recognition with kernel feature alignment
abstract
This paper proposes a novel algorithm for extracting informative motion features in daily life action recognition based on support vector machine (SVM). The main advantage of the proposed method is not only to extract remarkable motion features, which fit into human intuition, but also to improve the performance of the recognition system. Concretely speaking, the main properties of the proposed method are 1) optimizing kernel parameters so as to minimize its generalization error, 2) extracting remarkable motion features in response to the sensitivity of the kernel function. Experimental result shows that the proposed algorithm improves the accuracy of the recognition system and enables human to identify informative motion features intuitively.
Taketoshi Mori, Masamichi Shimosaka, Tatsuya Harada, Tomomasa Sato
IROS3
2003 Robot imitation of human motion based on qualitative description from multiple measurement of human and environmental data
abstract
This paper proposes an imitation algorithm for a robot to acquire typical tasks from multiple measured data of human tasks in the daily life. The algorithm consists of the following procedures: 1) Firstly, the system measures multiple human's object-transferring tasks on a table. Then it calculates qualitative description from measured raw data of positions of human hand and object, as well as the force applied to the table. This description is then converted to probabilistic description. 2) Secondly, the system finds typical human tasks with the maximum likelihood from the probabilistic description. 3) Thirdly, the trajectory, which enables the robot to imitate the typical human task, is extracted. 4) Finally, the imitation task within limited force to the environment is generated from the trajectory by simulation and adaptation. The experimental execution of the generated trajectory proves the validity of the algorithm.
Tomomasa Sato, Yuichiro Genda, Hideyuki Kubotera, Taketoshi Mori, Tatsuya Harada
IROS5
2003 Human behavior logging support system utilizing pose/position sensors and behavior target sensors
abstract
This paper proposes a behavior log creation support system utilizing human pose/position sensors and behavior target sensors. The system is equipped with a human pose sensor and a staying room sensors as the pose/position sensors as a well as a voice sensor and a PC utilization history sensor to detect the target of the behavior. The human pose sensor classifies human behaviors into "standing", "sitting", and "walking". The staying room sensor records a name of the room where the user stays. The voice sensor detects the conversation which appears when the user is communicating with someone else. The PC utilization history sensor records not only whether a PC is used or not but also the names of the application software to serve as a detector of the behavior target during the computer work. The measured data from these sensors is displayed to the user to support creating the behavior log, i.e. to support the user to recall and to input contents and targets of his or her behaviors. The experiment of making use of the sensors and creating the behavior log proved that the recorded events of the behavior log per all events improve from 60% with no support to more than 90% with support of the system. The result quantitatively shows the capability of human behavior logging support of the system.
Tomomasa Sato, Satoru Itoh, Satoshi Otani, Tatsuya Harada, Taketoshi Mori
IROS4
2002 Estimation of Bed-Ridden Human's Gross and Slight Movement Based on Pressure Sensors Distribution Bed
abstract
In this paper, we developed a bed-ridden human's body movement unrestraint estimation system by using a pressure sensors distribution bed. We classified body movements into gross and slight movements. In order to estimate both gross and slight movements, we realized the distinction methods between a human and an object and between sitting and lying status. We also realized the estimation methods of a posture, an articular movement, the respiration and the pulse. By integrating these methods to complement each other, a bed-ridden human's body movement from gross to slight movements can be estimated totally.
Tatsuya Harada, Tomomasa Sato, Taketoshi Mori
ICRA1
2001 Pressure Distribution Image Based Human Motion Tracking System Using Skeleton and Surface Integration Model
abstract
A lying person's motion tracking system by using a pressure distribution image and a full body model is proposed. The full body model consists of a skeleton and a surface model to cope with a variety of body shapes. BVH files are used as the skeleton model that describes a hierarchy of joints and links. Wavefront object files are used as the surface model that describes geometry of the surface. The bed has 210 pressure sensors that are under the mattress. It can measure a pressure distribution image of a lying person. The lying person's motion is tracked by considering potential energy, momentum and a difference between the measured pressure distribution image and a pressure distribution image that is calculated by the full body model. Experimental results reveal that the realized system can track not only horizontal motions such as opening and closing legs but also vertical motions such as raising the upper body.
Tatsuya Harada, Tomomasa Sato, Taketoshi Mori
ICRA1
2000 Infant Behavior Recognition System Based on Pressure Distribution Image
abstract
The authors developed a novel infant behavior recognition system based on a pressure distribution image. The system can recognize an infant's status (quiet, moving and crying), posture, body parts' positions and movement unrestrainedly. It can, in recognizing the behavior, cope with the infant's rapid growth and unique physique. The algorithm of the infant behavior recognition system is summarized as follows. 1) First, the system measures the pressure distribution image with 384 pressure sensors distributed in the bed. 2) The authors propose "activity score"; this is calculated by using the measured pressure distribution image and indicates kinetic energy of the infant's activity. Based on the activity score, the system decides the infant's status. 3) If the infant is quiet, the system estimates the infant's physique. 4) Based on the estimated physique, the system recognizes the infant's posture and body part movement. Experimental results reveal that the system successfully recognizes infants' status (quiet, moving and crying), posture, body parts position and movements.
Tatsuya Harada, Akihiko Saito, Tomomasa Sato, Taketoshi Mori
ICRA1
2000 Sensor pillow system: monitoring respiration and body movement in sleep
abstract
This paper presents "Sensor Pillow System" to measure physiological parameters in sleep without restraint to a human. The system consists of an array of pressure sensors under the pillow, a one-chip microcomputer to digitize and transmit the pressure data to a desktop computer; and the computer to count respirations and turns in sleep. This paper also presents a simple motion model which explains the change of the head pressure distribution accompanied with respiration. Based on this model, respiration count algorithms is proposed. The effectiveness of this system is experimentally shown by comparing the number of respirations and turns counted by the sensor pillow system of a medical device and a video image.
Tatsuya Harada, Akiko Sakata, Taketoshi Mori, Tomomasa Sato
IROS1
1999 Body Parts Positions and Posture Estimation System Based on Pressure Distribution Image
abstract
We develop a body parts position and posture estimation system. This system consists of a pressure sensor distributed bed and body parts position and posture estimation software. The computer constructs many pressure distribution images based on the simple human models and accumulates these images to its memory, called the model-based pressure image templates. The pressure distribution image through is compared with model-based pressure image templates to find out the most matched pressure image template. The model-based pressure image templates contain body parts positions and joint angles information, so the body parts position contacting with the bed can easily be estimated by using the most matched model-based pressure image template. Finally, an estimated posture is displayed as a 3D computer graphics image. An experimental result reveals that the system not only can display the estimated lying human posture intuitively, but can also estimate the lying human's body parts position accurately where the body parts contacts with bed.
Tatsuya Harada, Taketoshi Mori, Yoshifumi Nishida, Tomohisa Yoshimi, Tomomasa Sato
ICRA1
1997 Contact interaction robot-communication between robot and human through contact behavior
abstract
This paper proposes a contact interaction robot (CIR) which utilizes contact behavior as the interaction means between a human and a robot. The CIR is a puppet robot designed so that the robot and the human touch each other. The psychological experiments are performed by utilizing a CIR equipped with pressure sensors on both sides of its neck and six servo motors in its neck, two arms and, two legs. The experimental results reveal that the CIR is able to moderate the painfulness perceived by the human as well as to bring a sense of relief.
Tomomasa Sato, Tatsuya Harada, Taketoshi Mori
IROS2