EDBT 2026 Demo / reviewers in the wild / expert
Yoshitaka Ushiku
dblp:24/8843
· DBLP profile ↗
56ranked-venue papers
6as first author
23since 2021 · last 2025
0000-0002-9014-1389ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 12 since 2021Systems, architecture and hardware · 7 · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CaptionSmiths: Flexibly Controlling Language Pattern in Image CaptioningabstractAn image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated captions is not easy due to two reasons: (i) existing models are not given the properties as a condition during training and (ii) existing models cannot smoothly transition its language pattern from one state to the other. Given this challenge, we propose a new approach, CaptionSmiths, to acquire a single captioning model that can handle diverse language patterns. First, our approach quantifies three properties of each caption, length, descriptiveness, and uniqueness of a word, as continuous scalar values, without human annotation. Given the values, we represent the conditioning via interpolation between two endpoint vectors corresponding to the extreme states, e.g., one for a very short caption and one for a very long caption. Empirical results demonstrate that the resulting model can smoothly change the properties of the output captions and show higher lexical alignment than baselines. For instance, CaptionSmiths reduces the error in controlling caption length by 506\% despite better lexical alignment. Code will be available on https://github.com/omron-sinicx/captionsmiths. Kuniaki Saito, Donghyun Kim 0006, Kwanyong Park, Atsushi Hashimoto 0001, Yoshitaka Ushiku |
ICCV | 5 |
| 2025 | AgroBench: Vision-Language Model Benchmark in Agriculture
Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi, Yoshitaka Ushiku |
ICCV | 5 |
| 2025 | Rethinking the role of frames for SE(3)-invariant crystal structure modelingabstractCrystal structure modeling with graph neural networks is essential for various applications in materials informatics, and capturing SE(3)-invariant geometric features is a fundamental requirement for these networks. A straightforward approach is to model with orientation-standardized structures through structure-aligned coordinate systems, or “frames.” However, unlike molecules, determining frames for crystal structures is challenging due to their infinite and highly symmetric nature. In particular, existing methods rely on a statically fixed frame for each structure, determined solely by its structural information, regardless of the task under consideration. Here, we rethink the role of frames, *questioning whether such simplistic alignment with the structure is sufficient*, and propose the concept of *dynamic frames*. While accommodating the infinite and symmetric nature of crystals, these frames provide each atom with a dynamic view of its local environment, focusing on actively interacting atoms. We demonstrate this concept by utilizing the attention mechanism in a recent transformer-based crystal encoder, resulting in a new architecture called **CrystalFramer**. Extensive experiments show that CrystalFramer outperforms conventional frames and existing crystal encoders in various crystal property prediction tasks. Yusei Ito, Tatsunori Taniai, Ryo Igarashi 0002, Yoshitaka Ushiku, Kanta Ono |
ICLR | 4 |
| 2025 | SCU-Hand: Soft Conical Universal Robotic Hand for Scooping Granular Media from Containers of Various SizesabstractAutomating small-scale experiments in materials science presents challenges due to the heterogeneous nature of experimental setups. This study introduces the SCU-Hand (Soft Conical Universal Robot Hand), a novel end-effector designed to automate the task of scooping powdered samples from various container sizes using a robotic arm. The SCU-Hand employs a flexible, conical structure that adapts to different container geometries through deformation, maintaining consistent contact without complex force sensing or machine learning-based control methods. Its reconfigurable mechanism allows for size adjustment, enabling efficient scooping from diverse container types. By combining soft robotics principles with a sheet-morphing design, our end-effector achieves high flexibility while retaining the necessary stiffness for effective powder manipulation. We detail the design principles, fabrication process, and experimental validation of the SCU-Hand. Experimental validation showed that the scooping capacity is about 20% higher than that of a commercial tool, with a scooping performance of more than 95% for containers of sizes between 67 mm to 110 mm. This research contributes to laboratory automation by offering a cost-effective, easily implementable solution for automating tasks such as materials synthesis and characterization processes. Tomoya Takahashi, Cristian C. Beltran-Hernandez, Yuki Kuroda, Kazutoshi Tanaka, Masashi Hamaya, Yoshitaka Ushiku |
ICRA | 6 |
| 2025 | Where is the answer? An empirical study of positional bias for parametric knowledge extraction in language modelabstractKuniaki Saito, Chen-Yu Lee, Kihyuk Sohn, Yoshitaka Ushiku. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kuniaki Saito, Chen-Yu Lee, Kihyuk Sohn, Yoshitaka Ushiku |
NAACL (Long Papers) | 4 |
| 2025 | Cooperative Design Optimization through Natural Language Interactionabstractintegrating system-led optimization methods with Large Language Models (LLMs), allowing designers to intervene in the optimization process and better understand the system's reasoning.Experimental results show that our method provides higher user agency than a system-led method and shows promising optimization performance compared to manual design.It also matches the performance of an existing cooperative method with lower cognitive load. Ryogo Niwa, Shigeo Yoshida, Yuki Koyama 0001, Yoshitaka Ushiku |
UIST | 4 |
| 2025 | Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Yoichi Sato 0001 |
WACV | 6 |
| 2024 | Learning 3D Point Cloud Registration as a Single Optimization Problem
Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Shusaku Sone, Yoshitaka Ushiku |
ACCV (9) | 6 |
| 2024 | SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
Shohei Tanaka, Yoshitaka Ushiku |
BMVC | 3 |
| 2024 | COM Kitchens: An Unedited Overhead-View Video Dataset as a Vision-Language Benchmark
Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto 0001, Jun Harashima, Leszek Rybicki, Yusuke Fukasawa, Yoshitaka Ushiku |
ECCV (65) | 7 |
| 2024 | PolarDB: Formula-Driven Dataset for Pre-Training Trajectory EncodersabstractFormula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bias and racial bias because it does not rely on real data as discussed in previous studies using fractals and polygons for pre-training image encoders. While FDSL has been proposed for pre-training image encoders, it has not been considered for temporal trajectory data. In this paper, we introduce PolarDB, the first formula-driven dataset for pre-training trajectory encoders with an application to fine-grained cutting-method recognition using hand trajectories. More specifically, we generate 270k trajectories for 432 categories on the basis of polar equations and use them to pre-train a Transformer-based trajectory encoder in an FDSL manner. In the experiments, we show that pre-training on PolarDB improves the accuracy of fine-grained cutting-method recognition on cooking videos of EPIC-KITCHEN and Ego4D datasets, where the pre-trained trajectory encoder is used as a plug-in module for a video recognition network. Sota Miyamoto, Takuma Yagi, Yuto Makimoto, Mahiro Ukai, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Nakamasa Inoue |
ICASSP | 5 |
| 2024 | Crystalformer: Infinitely Connected Attention for Periodic Structure EncodingabstractPredicting physical properties of materials from their crystal structures is a fundamental problem in materials science. In peripheral areas such as the prediction of molecular properties, fully connected attention networks have been shown to be successful. However, unlike these finite atom arrangements, crystal structures are infinitely repeating, periodic arrangements of atoms, whose fully connected attention results in *infinitely connected attention*. In this work, we show that this infinitely connected attention can lead to a computationally tractable formulation, interpreted as *neural potential summation*, that performs infinite interatomic potential summations in a deeply learned feature space. We then propose a simple yet effective Transformer-based encoder architecture for crystal structures called *Crystalformer*. Compared to an existing Transformer-based model, the proposed model requires only 29.4% of the number of parameters, with minimal modifications to the original Transformer architecture. Despite the architectural simplicity, the proposed method outperforms state-of-the-art methods for various property regression tasks on the Materials Project and JARVIS-DFT datasets. Tatsunori Taniai, Ryo Igarashi 0002, Naoya Chiba, Kotaro Saito, Yoshitaka Ushiku, Kanta Ono |
ICLR | 6 |
| 2024 | Vision-Language Interpreter for Robot Task PlanningabstractLarge language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to generate a problem description (PD), a machine-readable file used by the planners to find a plan. By generating PDs from language instruction and scene observation, we can drive symbolic planners in a language-guided framework. We propose a Vision-Language Interpreter (ViLaIn), a new framework that generates PDs using state-of-the-art LLM and vision-language models. ViLaIn can refine generated PDs via error message feedback from the symbolic planner. Our aim is to answer the question: How accurately can ViLaIn and the symbolic planner generate valid robot plans? To evaluate ViLaIn, we introduce a novel dataset called the problem description generation (ProDG) dataset. The framework is evaluated with four new evaluation metrics. Experimental results show that ViLaIn can generate syntactically correct problems with more than 99% accuracy and valid plans with more than 58% accuracy. Our code and dataset are available at https://github.com/omron-sinicx/ViLaIn. Keisuke Shirai, Cristian C. Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto 0001, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, Shinsuke Mori |
ICRA | 8 |
| 2024 | AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringabstractVisual question answering aims to provide responses to questions given visual input. Recently, visual programmatic models (VPMs), which generate programs to answer questions through large language models (LLMs), have attracted attention. However, they often require long input prompts to provide the LLM with sufficient API usage details to generate relevant code. To address this limitation, we propose AdaCoder, an adaptive prompt compression framework for VPMs. AdaCoder operates in two phases: a compression phase and an inference phase. In the compression phase, given a preprompt that describes all API definitions with example code snippets, a set of compressed preprompts is generated, each depending on a specific question type. In the inference phase, AdaCoder predicts the question type and chooses the appropriate corresponding compressed preprompt to generate code to answer the question. In experiments, we apply AdaCoder to ViperGPT and demonstrate that it reduces token length by 71.1%, while maintaining or even improving the performance of visual question answering. Mahiro Ukai, Shuhei Kurita, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Nakamasa Inoue |
ACM Multimedia | 4 |
| 2023 | Robotic Powder Grinding with Audio-Visual Feedback for Laboratory Automation in Materials ScienceabstractThis study focuses on the powder grinding process, which is a necessary step for material synthesis in materials science experiments. In material science, powder grinding is a time-consuming process that is typically executed by hand, as commercial grinding machines are unsuitable for samples of small size. Robotic powder grinding would solve this problem, but it is a challenging task for robots, as it requires observing the powder state and generating appropriate motions. Our previous study proposed a robotic powder grinding system using visual feedback. Although visual feedback is helpful for observing the powder distribution, the particle size during the grinding process remains invisible, leading to suboptimal robot actions. In some cases, the robot chose to gather the powder even though continuing to grind instead would have produced finer powder. In this paper, we present a multi-modal robotic grinding system that utilizes both audio and visual feedback. It makes use of the grinding sound which carries information about the grinding progress, as the particle size strongly affects the audio intensity. The audio feedback enables the robot to grind until the powder is sufficiently fine. In our experiments, the robot ground 80.5% of the powder to a particle size smaller than$250\ \mu\mathrm{m}$with audio and visual feedback and 68% without audio feedback, indicating that multi-modal feedback is an effective tool to produce finer powder. We conclude that the addition of audio feedback provides crucial information to the robot, allowing it to better understand the progress of the grinding process and make more optimal decisions. This robot system can be used to prepare samples in material science experiments and analyze the grinding process. Yusaku Nakajima, Masashi Hamaya, Kazutoshi Tanaka, Takafumi Hawai, Felix von Drigalski, Yasuo Takeichi, Yoshitaka Ushiku, Kanta Ono |
IROS | 7 |
| 2023 | Reference-based Dense Pose Estimation via Partial 3D Point Cloud MatchingabstractInteracting with real-world objects is one of the fundamental tasks in multimedia. Despite its importance, existing object pose estimation targets only rigid objects. This demonstration proposes a novel application for non-rigid object pose estimation. Inspired by human dense pose estimation, we represent a pose of a non-rigid object as an indexed point cloud, where each index corresponds to that in a template. The correspondence is identified by a machine-learning-based 3D point cloud matching. Finding correspondence to the template point cloud enables a dense pose estimation with no object-specific learning processes. In the demonstration, we visualize the correspondence of points in observed depth images and the template. We also provide a demonstration of template point cloud reconstruction. Through these systems, onsite visitors can test our system with objects brought by themselves and have an experience with a state-of-the-art 3D point cloud matching method as well as this novel task. Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Yoshitaka Ushiku |
ACM Multimedia | 4 |
| 2023 | State-aware video procedural captioning
Taichi Nishimura, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Hirotaka Kameko, Shinsuke Mori |
Multim. Tools Appl. | 3 |
| 2022 | Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe FlowsabstractWe present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is represented as a recipe flow graph. We developed a web interface to reduce human annotation costs. The dataset allows us to try various applications, including multimodal information retrieval. Keisuke Shirai, Atsushi Hashimoto 0001, Taichi Nishimura, Hirotaka Kameko, Shuhei Kurita, Yoshitaka Ushiku, Shinsuke Mori |
COLING | 6 |
| 2022 | Robotic Powder Grinding with a Soft Jig for Laboratory Automation in Material ScienceabstractGrinding materials into a fine powder is a time-consuming task in material science that is generally performed by hand, as current automated grinding machines might not be suitable for preparing small-sized samples. This study presents a robotic powder grinding system for laboratory automation in material science applications that observe the powder's state to improve the grinding outcome. We developed a soft jig consisting of off-the-shelf gel materials and 3D-printed parts, which can be used with any robot arm to perform powder grinding. The jig's physical softness allows for safe grinding without force sensing. In addition, we developed a visual feedback system that observes the powder distribution and decides where to grind and when to gather. The results showed that our system could grind 79 percent of the powder to a particle size smaller than 200 μm by using the soft jig and visual feedback. This ratio was 57% when using only the soft jig without feedback. Our system can be used immediately in laboratories to alleviate the workload of researchers. Yusaku Nakajima, Masashi Hamaya, Takafumi Hawai, Felix von Drigalski, Kazutoshi Tanaka, Yoshitaka Ushiku, Kanta Ono |
IROS | 7 |
| 2021 | Divergence Optimization for Noisy Universal Domain AdaptationabstractUniversal domain adaptation (UniDA) has been proposed to transfer knowledge learned from a label-rich source domain to a label-scarce target domain without any constraints on the label sets. In practice, however, it is difficult to obtain a large amount of perfectly clean labeled data in a source domain with limited resources. Existing UniDA methods rely on source samples with correct annotations, which greatly limits their application in the real world. Hence, we consider a new realistic setting called Noisy UniDA, in which classifiers are trained with noisy labeled data from the source domain and unlabeled data with an unknown class distribution from the target domain. This paper introduces a two-head convolutional neural network framework to solve all problems simultaneously. Our network consists of one common feature generator and two classifiers with different decision boundaries. By optimizing the divergence between the two classifiers’ outputs, we can detect noisy source samples, find "unknown" classes in the target domain, and align the distribution of the source and target domains. In an extensive evaluation of different domain adaptation settings, the proposed method outperformed existing methods by a large margin in most settings. Qing Yu 0013, Atsushi Hashimoto 0001, Yoshitaka Ushiku |
CVPR | 3 |
| 2021 | Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image CaptioningabstractUkyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto, Taro Watanabe, Yuji Matsumoto. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Taro Watanabe, Yuji Matsumoto 0001 |
EACL | 2 |
| 2021 | State-aware Video Procedural CaptioningabstractVideo procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material list. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition. The code and dataset are available from https://github.com/misogil0116/svpc Taichi Nishimura, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Hirotaka Kameko, Shinsuke Mori |
ACM Multimedia | 3 |
| 2021 | Retransmission Edge Computing System Conducting Adaptive Image Compression Based on Image Recognition AccuracyabstractThis paper proposes a retransmission control system based on image recognition accuracy as a traffic reduction technique for improving network bandwidth usage efficiency in image recognition service using wireless edge computing. For traffic reduction, image compression is useful. However, it is known to deteriorate the recognition accuracy. Our proposed system is applied to guarantee this deterioration. By retransmitting images according to the recognition accuracy, we aim to reduce traffic and to guarantee recognition accuracy. When compressing the image, PSNR is calculated to adaptively change the ratio of the image compression, to maintain the image quality. This paper uses down-sampler as an image compression method and demonstrates the effectiveness of the proposed system through a wireless network simulation. We confirm that our proposed retransmission system can reduce the network traffic congestion, and guarantee the recognition accuracy as same as the conventional method. Mutsuki Nakahara, Daisuke Hisano, Mai Nishimura, Yoshitaka Ushiku, Kazuki Maruta, Yu Nakayama |
VTC Fall | 4 |
| 2020 | Visual Grounding Annotation of Recipe Flow GraphabstractIn this paper, we provide a dataset that gives visual grounding annotations to recipe flow graphs. A recipe flow graph is a representation of the cooking workflow, which is designed with the aim of understanding the workflow from natural language processing. Such a workflow will increase its value when grounded to real-world activities, and visual grounding is a way to do so. Visual grounding is provided as bounding boxes to image sequences of recipes, and each bounding box is linked to an element of the workflow. Because the workflows are also linked to the text, this annotation gives visual grounding with workflow’s contextual information between procedural text and visual observation in an indirect manner. We subsidiarily annotated two types of event attributes with each bounding box: “doing-the-action,” or “done-the-action”. As a result of the annotation, we got 2,300 bounding boxes in 272 flow graph recipes. Various experiments showed that the proposed dataset enables us to estimate contextual information described in recipe flow graphs from an image sequence. Taichi Nishimura, Suzushi Tomori, Hayato Hashimoto, Atsushi Hashimoto 0001, Yoko Yamakata, Jun Harashima, Yoshitaka Ushiku, Shinsuke Mori |
LREC | 7 |
| 2019 | Estimating the Causal Effect from Partially Observed Time SeriesabstractMany real-world systems involve interacting time series. The ability to detect causal dependencies between system components from observed time series of their outputs is essential for understanding system behavior. The quantification of causal influences between time series is based on the definition of some causality measure. Partial Canonical Correlation Analysis (Partial CCA) and its extensions are examples of methods used for robustly estimating the causal relationships between two multidimensional time series even when the time series are short. These methods assume that the input data are complete and have no missing values. However, real-world data often contain missing values. It is therefore crucial to estimate the causality measure robustly even when the input time series is incomplete. Treating this problem as a semi-supervised learning problem, we propose a novel semi-supervised extension of probabilistic Partial CCA called semi-Bayesian Partial CCA. Our method exploits the information in samples with missing values to prevent the overfitting of parameter estimation even when there are few complete samples. Experiments based on synthesized and real data demonstrate the ability of the proposed method to estimate causal relationships more correctly than existing methods when the data contain missing values, the dimensionality is large, and the number of samples is small. Akane Iseki, Yusuke Mukuta, Yoshitaka Ushiku, Tatsuya Harada |
AAAI | 3 |
| 2019 | Class-Distinct and Class-Mutual Image Generation with GANs
Takuhiro Kaneko, Yoshitaka Ushiku, Tatsuya Harada |
BMVC | 2 |
| 2019 | Label-Noise Robust Generative Adversarial NetworksabstractGenerative adversarial networks (GANs) are a framework that learns a generative distribution through adversarial training. Recently, their class conditional extensions (e.g., conditional GAN (cGAN) and auxiliary classifier GAN (AC-GAN)) have attracted much attention owing to their ability to learn the disentangled representations and to improve the training stability. However, their training requires the availability of large-scale accurate class-labeled data, which are often laborious or impractical to collect in a real-world scenario. To remedy this, we propose a novel family of GANs called label-noise robust GANs (rGANs), which, by incorporating a noise transition model, can learn a clean label conditional generative distribution even when training labels are noisy. In particular, we propose two variants: rAC-GAN, which is a bridging model between AC-GAN and the label-noise robust classification model, and rcGAN, which is an extension of cGAN and solves this problem with no reliance on any classifier. In addition to providing the theoretical background, we demonstrate the effectiveness of our models through extensive experiments using diverse GAN configurations, various noise settings, and multiple evaluation metrics (in which we tested 402 conditions in total). Takuhiro Kaneko, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 2 |
| 2019 | Strong-Weak Distribution Alignment for Adaptive Object DetectionabstractWe propose an approach for unsupervised adaptation of object detectors from label-rich to label-poor domains which can significantly reduce annotation costs associated with detection. Recently, approaches that align distributions of source and target images using an adversarial loss have been proven effective for adapting object classifiers. However, for object detection, fully matching the entire distributions of source and target images to each other at the global image level may fail, as domains could have distinct scene layouts and different combinations of objects. On the other hand, strong matching of local features such as texture and color makes sense, as it does not change category level semantics. This motivates us to propose a novel method for detector adaptation based on strong local alignment and weak global alignment. Our key contribution is the weak alignment model, which focuses the adversarial alignment loss on images that are globally similar and puts less emphasis on aligning images that are globally dissimilar. Additionally, we design the strong domain alignment model to only look at local receptive fields of the feature map. We empirically verify the effectiveness of our method on four datasets comprising both large and small domain shifts. Our code is available at https://github.com/VisionLearningGroup/DA_Detection. Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, Kate Saenko |
CVPR | 2 |
| 2019 | Generating Easy-to-Understand Referring Expressions for Target IdentificationsabstractThis paper addresses the generation of referring expressions that not only refer to objects correctly but also let humans find them quickly. As a target becomes relatively less salient, identifying referred objects itself becomes more difficult. However, the existing studies regarded all sentences that refer to objects correctly as equally good, ignoring whether they are easily understood by humans. If the target is not salient, humans utilize relationships with the salient contexts around it to help listeners to comprehend it better. To derive this information from human annotations, our model is designed to extract information from the target and from the environment. Moreover, we regard that sentences that are easily understood are those that are comprehended correctly and quickly by humans. We optimized this by using the time required to locate the referred objects by humans and their accuracies. To evaluate our system, we created a new referring expression dataset whose images were acquired from Grand Theft Auto V (GTA V), limiting targets to persons. Experimental results show the effectiveness of our approach. Our code and dataset are available at https://github.com/mikittt/easy-to-understand-REG. Mikihiro Tanaka, Takayuki Itamochi, Kenichi Narioka, Ikuro Sato, Yoshitaka Ushiku, Tatsuya Harada |
ICCV | 5 |
| 2019 | Pose Graph optimization for Unsupervised Monocular Visual OdometryabstractUnsupervised Learning based monocular visual odometry (VO) has lately drawn significant attention for its potential in label-free leaning ability and robustness to camera parameters and environmental variations. However, partially due to the lack of drift correction technique, these methods are still by far less accurate than geometric approaches for large-scale odometry estimation. In this paper, we propose to leverage graph optimization and loop closure detection to overcome limitations of unsupervised learning based monocular visual odometry. To this end, we propose a hybrid VO system which combines an unsupervised monocular VO called NeuralBundler with a pose graph optimization back-end. NeuralBundler is a neural network architecture that uses temporal and spatial photometric loss as main supervision and generates a windowed pose graph consists of multi-view 6DoF constraints. We propose a novel pose cycle consistency loss to relieve the tensions in the windowed pose graph, leading to improved performance and robustness. In the back-end, a global pose graph is built from local and loop 6DoF constraints estimated by NeuralBundler, and is optimized over SE(3). Empirical evaluation on the KITTI odometry dataset demonstrates that 1) NeuralBundler achieves state-of-the-art performance on unsupervised monocular VO estimation, and 2) our whole approach can achieve efficient loop closing and show favorable overall translational accuracy compared to established monocular SLAM systems. Yang Li 0143, Yoshitaka Ushiku, Tatsuya Harada |
ICRA | 2 |
| 2019 | How narratives move your mind: A corpus of shared-character stories for connecting emotional flow and interestingness
Yusuke Mori 0001, Hiroaki Yamane, Yoshitaka Ushiku, Tatsuya Harada |
Inf. Process. Manag. | 3 |
| 2018 | Alternating Circulant Random Features for Semigroup Kernels
Yusuke Mukuta, Yoshitaka Ushiku, Tatsuya Harada |
AAAI | 2 |
| 2018 | Hierarchical Video Generation From Orthogonal Information: Optical Flow and TextureabstractLearning to represent and generate videos from unlabeled data is a very challenging problem. To generate realistic videos, it is important not only to ensure that the appearance of each frame is real, but also to ensure the plausibility of a video motion and consistency of a video appearance in the time direction. The process of video generation should be divided according to these intrinsic difficulties. In this study, we focus on the motion and appearance information as two important orthogonal components of a video, and propose Flow-and-Texture-Generative Adversarial Networks (FTGAN) consisting of FlowGAN and TextureGAN. In order to avoid a huge annotation cost, we have to explore a way to learn from unlabeled data. Thus, we employ optical flow as motion information to generate videos. FlowGAN generates optical flow, which contains only the edge and motion of the videos to be begerated. On the other hand, TextureGAN specializes in giving a texture to optical flow generated by FlowGAN. This hierarchical approach brings more realistic videos with plausible motion and appearance consistency. Our experiments show that our model generates more plausible motion videos and also achieves significantly improved performance for unsupervised action classification in comparison to previous GAN works. In addition, because our model generates videos from two independent information, our model can generate new combinations of motion and attribute that are not seen in training data, such as a video in which a person is doing sit-up in a baseball ground. Katsunori Ohnishi, Shohei Yamamoto, Yoshitaka Ushiku, Tatsuya Harada |
AAAI | 3 |
| 2018 | Viewpoint-Aware Video SummarizationabstractThis paper introduces a novel variant of video summarization, namely building a summary that depends on the particular aspect of a video the viewer focuses on. We refer to this as viewpoint. To infer what the desired viewpoint may be, we assume that several other videos are available, especially groups of videos, e.g., as folders on a person's phone or laptop. The semantic similarity between videos in a group vs. the dissimilarity between groups is used to produce viewpoint-specific summaries. For considering similarity as well as avoiding redundancy, output summary should be (A) diverse, (B) representative of videos in the same group, and (C) discriminative against videos in the different groups. To satisfy these requirements (A)-(C) simultaneously, we proposed a novel video summarization method from multiple groups of videos. Inspired by Fisher's discriminant criteria, it selects summary by optimizing the combination of three terms (a) inner-summary, (b) inner-group, and (c) between-group variances defined on the feature representation of summary, which can simply represent (A)-(C). Moreover, we developed a novel dataset to investigate how well the generated summary reflects the underlying viewpoint. Quantitative and qualitative experiments conducted on the dataset demonstrate the effectiveness of proposed method. Atsushi Kanehira, Luc Van Gool, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 3 |
| 2018 | Neural 3D Mesh RendererabstractFor modeling the 3D world behind 2D images, which 3D representation is most appropriate? A polygon mesh is a promising candidate for its compactness and geometric properties. However, it is not straightforward to model a polygon mesh from 2D images using neural networks because the conversion from a mesh to an image, or rendering, involves a discrete operation called rasterization, which prevents back-propagation. Therefore, in this work, we propose an approximate gradient for rasterization that enables the integration of rendering into neural networks. Using this renderer, we perform single-image 3D mesh reconstruction with silhouette image supervision and our system outperforms the existing voxel-based approach. Additionally, we perform gradient-based 3D mesh editing operations, such as 2D-to-3D style transfer and 3D DeepDream, with 2D supervision for the first time. These applications demonstrate the potential of the integration of a mesh renderer into neural networks and the effectiveness of our proposed renderer. Hiroharu Kato, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 2 |
| 2018 | Maximum Classifier Discrepancy for Unsupervised Domain AdaptationabstractIn this work, we present a method for unsupervised domain adaptation. Many adversarial learning methods train domain classifier networks to distinguish the features as either a source or target and train a feature generator network to mimic the discriminator. Two problems exist with these methods. First, the domain classifier only tries to distinguish the features as a source or target and thus does not consider task-specific decision boundaries between classes. Therefore, a trained generator can generate ambiguous features near class boundaries. Second, these methods aim to completely match the feature distributions between different domains, which is difficult because of each domain's characteristics. To solve these problems, we introduce a new approach that attempts to align distributions of source and target by utilizing the task-specific decision boundaries. We propose to maximize the discrepancy between two classifiers' outputs to detect target samples that are far from the support of the source. A feature generator learns to generate target features near the support to minimize the discrepancy. Our method outperforms other methods on several datasets of image classification and semantic segmentation. The codes are available at https://github.com/mil-tokyo/MCD_DA. Kuniaki Saito, Kohei Watanabe, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 3 |
| 2018 | Customized Image Narrative Generation via Interactive Visual Question Generation and AnsweringabstractImage description task has been invariably examined in a static manner with qualitative presumptions held to be universally applicable, regardless of the scope or target of the description. In practice, however, different viewers may pay attention to different aspects of the image, and yield different descriptions or interpretations under various contexts. Such diversity in perspectives is difficult to derive with conventional image description techniques. In this paper, we propose a customized image narrative generation task, in which the users are interactively engaged in the generation process by providing answers to the questions. We further attempt to learn the user's interest via repeating such interactive stages, and to automatically reflect the interest in descriptions for new images. Experimental results demonstrate that our model can generate a variety of descriptions from single image that cover a wider range of topics than conventional models, while being customizable to the target user of interaction. Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 2 |
| 2018 | Between-Class Learning for Image ClassificationabstractIn this paper, we propose a novel learning method for image classification called Between-Class learning (BC learning)1. We generate between-class images by mixing two images belonging to different classes with a random ratio. We then input the mixed image to the model and train the model to output the mixing ratio. BC learning has the ability to impose constraints on the shape of the feature distributions, and thus the generalization ability is improved. BC learning is originally a method developed for sounds, which can be digitally mixed. Mixing two image data does not appear to make sense; however, we argue that because convolutional neural networks have an aspect of treating input data as waveforms, what works on sounds must also work on images. First, we propose a simple mixing method using internal divisions, which surprisingly proves to significantly improve performance. Second, we propose a mixing method that treats the images as waveforms, which leads to a further improvement in performance. As a result, we achieved 19.4% and 2.26% top-1 errors on ImageNet-1K and CIFAR10, respectively. Yuji Tokozume, Yoshitaka Ushiku, Tatsuya Harada |
CVPR | 2 |
| 2018 | Open Set Domain Adaptation by Backpropagation
Kuniaki Saito, Shohei Yamamoto, Yoshitaka Ushiku, Tatsuya Harada |
ECCV (5) | 3 |
| 2018 | Visual Question Generation for Class Acquisition of Unknown Objects
Kohei Uehara, Antonio Tejero-de-Pablos, Yoshitaka Ushiku, Tatsuya Harada |
ECCV (12) | 3 |
| 2018 | Adversarial Dropout Regularization
Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, Kate Saenko |
ICLR (Poster) | 2 |
| 2018 | Learning from Between-class Examples for Deep Sound Recognition
Yuji Tokozume, Yoshitaka Ushiku, Tatsuya Harada |
ICLR (Poster) | 2 |
| 2017 | Spatio-Temporal Person Retrieval via Natural Language QueriesabstractIn this paper, we address the problem of spatio-temporal person retrieval from videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query. For this problem, we introduce a novel dataset consisting of videos containing people annotated with bounding boxes for each second and with five natural language descriptions. To retrieve the tube of the person described by a given natural language query, we design a model that combines methods for spatio-temporal human detection and multimodal retrieval. We conduct comprehensive experiments to compare a variety of tube and text representations and multimodal retrieval methods, and present a strong baseline in this task as well as demonstrate the efficacy of our tube representation and multimodal feature embedding technique. Finally, we demonstrate the versatility of our model by applying it to two other important tasks. Masataka Yamaguchi, Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada |
ICCV | 3 |
| 2017 | DualNet: Domain-invariant network for visual question answeringabstractVisual question answering (VQA) tasks use two types of images: abstract (illustrations) and real. Domain-specific differences exist between the two types of images with respect to “objectness,” “texture,” and “color.” Therefore, achieving similar performance by applying methods developed for real images to abstract images, and vice versa, is difficult. This is a critical problem in VQA, because image features are crucial clues for correctly answering the questions about the images. However, an effective, domain-invariant method can provide insight into the high-level reasoning required for VQA. We thus propose a method called DualNet that demonstrates performance that is invariant to the differences in real and abstract scene domains. Experimental results show that DualNet outperforms state-of-the-art methods, especially for the abstract images category. Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada |
ICME | 3 |
| 2017 | Asymmetric Tri-training for Unsupervised Domain AdaptationabstractIt is important to apply models trained on a large number of labeled samples to different domains because collecting many labeled samples in various domains is expensive. To learn discriminative representations for the target domain, we assume that artificially labeling the target samples can result in a good representation. Tri-training leverages three classifiers equally to provide pseudo-labels to unlabeled samples; however, the method does not assume labeling samples generated from a different domain. In this paper, we propose the use of an asymmetric tri-training method for unsupervised domain adaptation, where we assign pseudo-labels to unlabeled samples and train the neural networks as if they are true labels. In our work, we use three networks asymmetrically, and by asymmetric, we mean that two networks are used to label unlabeled target samples, and one network is trained by the pseudo-labeled samples to obtain target-discriminative representations. Our proposed method was shown to achieve a state-of-the-art performance on the benchmark digit recognition datasets for domain adaptation. Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada |
ICML | 2 |
| 2017 | MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenesabstractThis work addresses the semantic segmentation of images of street scenes for autonomous vehicles based on a new RGB-Thermal dataset, which is also introduced in this paper. An increasing interest in self-driving vehicles has brought the adaptation of semantic segmentation to self-driving systems. However, recent research relating to semantic segmentation is mainly based on RGB images acquired during times of poor visibility at night and under adverse weather conditions. Furthermore, most of these methods only focused on improving performance while ignoring time consumption. The aforementioned problems prompted us to propose a new convolutional neural network architecture for multi-spectral image segmentation that enables the segmentation accuracy to be retained during real-time operation. We benchmarked our method by creating an RGB-Thermal dataset in which thermal and RGB images are combined. We showed that the segmentation accuracy was significantly increased by adding thermal infrared information. Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, Tatsuya Harada |
IROS | 4 |
| 2017 | WebDNN: Fastest DNN Execution Framework on Web BrowserabstractRecently, deep neural network (DNN) is drawing a lot of attention because of its applications. However, it requires a lot of computational resources and tremendous processes in order to setup an execution environment based on hardware acceleration such as GPGPU. Therefore, providing DNN applications to end-users is very hard. To solve this problem, we have developed an installation-free web browser-based DNN execution framework, WebDNN. WebDNN optimizes the trained DNN model to compress model data and accelerate the execution. It executes the DNN model with novel JavaScript API to achieve zero-overhead execution. Empirical evaluations show that it achieves more than two-hundred times the unusual acceleration. WebDNN is an open source framework and you can download it from https://github.com/mil-tokyo/webdnn. Masatoshi Hidaka, Yuichiro Kikura, Yoshitaka Ushiku, Tatsuya Harada |
ACM Multimedia | 3 |
| 2016 | Image Captioning with Sentiment Terms via Weakly-Supervised Sentiment Dataset
Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada |
BMVC | 2 |
| 2015 | Common Subspace for Model and Similarity: Phrase Learning for Caption Generation from ImagesabstractGenerating captions to describe images is a fundamental problem that combines computer vision and natural language processing. Recent works focus on descriptive phrases, such as "a white dog" to explain the visual composites of an input image. The phrases can not only express objects, attributes, events, and their relations but can also reduce visual complexity. A caption for an input image can be generated by connecting estimated phrases using a grammar model. However, because phrases are combinations of various words, the number of phrases is much larger than the number of single words. Consequently, the accuracy of phrase estimation suffers from too few training samples per phrase. In this paper, we propose a novel phrase-learning method: Common Subspace for Model and Similarity (CoSMoS). In order to overcome the shortage of training samples, CoSMoS obtains a subspace in which (a) all feature vectors associated with the same phrase are mapped as mutually close, (b) classifiers for each phrase are learned, and (c) training samples are shared among co-occurring phrases. Experimental results demonstrate that our system is more accurate than those in earlier work and that the accuracy increases when the dataset from the web increases. Yoshitaka Ushiku, Masataka Yamaguchi, Yusuke Mukuta, Tatsuya Harada |
ICCV | 1 |
| 2014 | Three Guidelines of Online Learning for Large-Scale Visual RecognitionabstractIn this paper, we would like to evaluate online learning algorithms for large-scale visual recognition using state-of-the-art features which are preselected and held fixed. Today, combinations of high-dimensional features and linear classifiers are widely used for large-scale visual recognition. Numerous so-called mid-level features have been developed and mutually compared on an experimental basis. Although various learning methods for linear classification have also been proposed in the machine learning and natural language processing literature, they have rarely been evaluated for visual recognition. Therefore, we give guidelines via investigations of state-of-the-art online learning methods of linear classifiers. Many methods have been evaluated using toy data and natural language processing problems such as document classification. Consequently, we gave those methods a unified interpretation from the viewpoint of visual recognition. Results of controlled comparisons indicate three guidelines that might change the pipeline for visual recognition. Yoshitaka Ushiku, Masatoshi Hidaka, Tatsuya Harada |
CVPR | 1 |
| 2014 | Hard negative classes for multiple object detectionabstractWe propose an efficient method to train multiple object detectors simultaneously using a large scale image dataset. The one-vs-all approach that optimizes the boundary between positive samples from a target class and negative samples from the others has been the most standard approach for object detection. However, because this approach trains each object detector independently, the scores are not balanced between object classes. The proposed method combines ideas derived from both detection and classification in order to balance the scores across all object classes. We optimized the boundary between target classes and their “hard negative” samples, just as in detection, while simultaneously balancing the detector scores across object classes, as done in multi-class classification. We evaluated the performances on multi-class object detection using a subset of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2011 dataset and showed our method outperformed a de facto standard method. Asako Kanezaki, Sho Inaba, Yoshitaka Ushiku, Yuya Yamashita, Hiroshi Muraoka, Yasuo Kuniyoshi, Tatsuya Harada |
ICRA | 3 |
| 2012 | Efficient image annotation for automatic sentence generationabstractSentence generation from images is an ultimate goal of image recognition. In this paper, we attack a novel problem, the "multi-keyphrase problem", to address this goal. We hypothesize that image contents can be described with multi-keyphrases, and that a natural sentence can be generated by connecting multi-keyphrases with an experimental grammar model. Existing methods require semantic knowledge such as labels of an object, action, or scene. Using these methods, we must strive to prepare a highly organized dataset. Therefore, we propose a novel online learning method for multi-keyphrase estimation. The proposed framework, although simple and scalable, can generate sentences from images with no semantic knowledge. Moreover, the proposed method for multi-keyphrase estimation is applicable to image annotation, and it achieves state-of-the-art performance. Our experiment using only images and texts demonstrates that the proposed framework is useful for sentence generation from images. Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi |
ACM Multimedia | 1 |
| 2011 | Discriminative spatial pyramidabstractSpatial Pyramid Representation (SPR) is a widely used method for embedding both global and local spatial information into a feature, and it shows good performance in terms of generic image recognition. In SPR, the image is divided into a sequence of increasingly finer grids on each pyramid level. Features are extracted from all of the grid cells and are concatenated to form one huge feature vector. As a result, expensive computational costs are required for both learning and testing. Moreover, because the strategy for partitioning the image at each pyramid level is designed by hand, there is weak theoretical evidence of the appropriate partitioning strategy for good categorization. In this paper, we propose discriminative SPR, which is a new representation that forms the image feature as a weighted sum of semi-local features over all pyramid levels. The weights are automatically selected to maximize a discriminative power. The resulting feature is compact and preserves high discriminative power, even in low dimension. Furthermore, the discriminative SPR can suggest the distinctive cells and the pyramid levels simultaneously by observing the optimal weights generated from the fine grid cells. Tatsuya Harada, Yoshitaka Ushiku, Yuya Yamashita, Yasuo Kuniyoshi |
CVPR | 2 |
| 2011 | Understanding images with natural sentencesabstractWe propose a novel system which generates sentential captions for general images. For people to use numerous images effectively on the web, technologies must be able to explain image contents and must be capable of searching for data that users need. Moreover, images must be described with natural sentences based not only on the names of objects contained in an image but also on their mutual relations. The proposed system uses general images and captions available on the web as training data to generate captions for new images. Furthermore, because the learning cost is independent from the amount of data, the system has scalability, which makes it useful with large-scale data. Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi |
ACM Multimedia | 1 |
| 2011 | Automatic sentence generation from imagesabstractFor the overwhelming amounts of multimedia used on the Web, methods of search and understanding with sentences are necessary. Representing the contents not only using labels but also using sentences including labels' relations enables users to search with a story and to understand multimedia deeply. However, few existing works describe such sentences because obtaining objects' relations and grammar is difficult. We specifically examine captions of images that are similar to an input image. They are expected to explain the input image to some degree. Therefore, we propose a novel approach to generate a sentential caption for the input image by summarizing those captions. Our experiment using a dataset consisting of images and text demonstrates that the proposed method can generate sentential captions. Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi |
ACM Multimedia | 1 |
| 2010 | Improving image similarity measures for image browsing and retrieval through latent space learning between images and long textsabstractThe amount of multimedia data on personal devices and the Web is increasing daily. Image browsing and retrieval systems in a low-dimensional space have been widely studied to manage and view large numbers of images. It is essential for such systems to exploit an efficient similarity measure of the images when searching for them. Existing methods use the distance in a low-level image feature space as the similarity measure, and therefore, images with different content may be treated as similar images. In this paper, we propose a novel method to improve the similarity measures for images by considering the text surrounding the images. If there is text describing the images, similarities can be measured more effectively by taking into account the text streams. The proposed method improves the image similarity measures based on the latent semantics obtained from the combination of image and text. It should be noted that the text does not need to be clear tags; indeed, any generic Web text is applicable. Moreover, our method can effectively improve the similarities even if only a small portion of the images include textual descriptions. Additionally, the proposed method is scalable as it has linear computational complexity based on the number of images. In the experiments, we compare our method with previous methods using an original dataset in which a portion of the images are annotated by long text. We show that the proposed method can retrieve semantically similar images more precisely than existing methods. Yoshitaka Ushiku, Tatsuya Harada, Yasuo Kuniyoshi |
ICIP | 1 |