Hideki Nakayama

dblp:09/1592 · DBLP profile ↗
← Back
78ranked-venue papers
6as first author
32since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 21 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
abstract
Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) mod- els using residual vector quantization (RVQ). However, long-form speech synthe- sis remains a significant challenge due to the high frame rate, which increases the length of audio tokens and makes it difficult for autoregressive language models to generate audio tokens for even a minute of speech. To address this challenge, this paper introduces two novel post-training approaches: 1) Multi-Resolution Re- quantization (MReQ) and 2) HALL-E. MReQ is a framework to reduce the frame rate of pre-trained NAC models. Specifically, it incorporates multi-resolution residual vector quantization (MRVQ) module that hierarchically reorganizes dis- crete audio tokens through teacher-student distillation. HALL-E is an LLM-based TTS model designed to predict hierarchical tokens of MReQ. Specifically, it incor- porates the technique of using MRVQ sub-modules and continues training from a pre-trained LLM-based TTS model. Furthermore, to promote TTS research, we create MinutesSpeech, a new benchmark dataset consisting of 40k hours of filtered speech data for training and evaluating speech synthesis ranging from 3s up to 180s. In experiments, we demonstrated the effectiveness of our approaches by ap- plying our post-training framework to VALL-E. We achieved the frame rate down to as low as 8 Hz, enabling the stable minitue-long speech synthesis in a single inference step. Audio samples, dataset, codes and pre-trained models are available at https://yutonishimura-v2.github.io/HALL-E_DEMO.
Yuto Nishimura, Takumi Hirose, Masanari Ohi, Hideki Nakayama, Nakamasa Inoue
ICLR4
2025 LAVA Grand Challenge 2025: Benchmarking Japanese-English Document Understanding with Large Vision-Language Models
abstract
The advent of Large Vision-Language Models (LVLMs) has demonstrated significant capabilities in multimodal understanding. However, their application to complex, multi-page documents, particularly in non-English languages like Japanese, remains a significant challenge due to the scarcity of suitable benchmarks. To address this gap, we organized the ''Large Vision---Language Model Learning and Applications (LAVA) Grand Challenge'' at ACM MultiMedia 2025. We present an overview of the competition. We designed a novel, challenging task: a 10-way multiple-choice Visual Question Answering (VQA) task on multi-page Japanese PDF documents. The task demands that models integrate information across multiple pages, text, and figures. We detail the dataset construction, including an annotation and filtering process designed to ensure questions are visually grounded and non-trivial. We also present the competition results, including an analysis of the leaderboard, and discuss the baseline performance of representative models. The LAVA Grand Challenge highlighted both the current capabilities and limitations of LVLMs in practical document understanding scenarios, thereby stimulating future research and providing a robust benchmark in this important domain.
Daichi Sato, Duc Minh Vo, Khan Md. Anwarus Salam, Hidenori Shoji, Yuma Matsuoka, Takara Taniguchi, Kaito Baba, Hideki Nakayama
ACM Multimedia8
2024 Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
abstract
Enhancing user engagement through personalization in conversational agents has gained significance, especially with the advent of large language models that generate fluent responses. Personalized dialogue generation, however, is multifaceted and varies in its definition – ranging from instilling a persona in the agent to capturing users’ explicit and implicit cues. This paper seeks to systemically survey the recent landscape of personalized dialogue generation, including the datasets employed, methodologies developed, and evaluation metrics applied. Covering 22 datasets, we highlight benchmark datasets and newer ones enriched with additional features. We further analyze 17 seminal works from top conferences between 2021-2023 and identify five distinct types of problems. We also shed light on recent progress by LLMs in personalized dialogue generation. Our evaluation section offers a comprehensive summary of assessment facets and metrics utilized in these works. In conclusion, we discuss prevailing challenges and envision prospect directions for future research in personalized dialogue generation.
Yi-Pei Chen 0001, Noriki Nishida, Hideki Nakayama, Yuji Matsumoto 0001
LREC/COLING3
2024 Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
abstract
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrieval-augmented image captioning method that prompts LLMs with object names retrieved from External Visual-name memory (EVCAP). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effort-lessly augment LLMs with retrieved object names by uti-lizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or retraining. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EV-CAP, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pretrained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
Jiaxuan Li 0004, Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
CVPR4
2024 LayoutFlow: Flow Matching for Layout Generation
Julian Jorge Andrade Guerreiro, Naoto Inoue, Kento Masui, Mayu Otani, Hideki Nakayama
ECCV (36)5
2024 A Compact Dynamic 3D Gaussian Representation for Real-Time Dynamic View Synthesis
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
ECCV (86)3
2024 Soft Curriculum for Learning Conditional GANs with Noisy-Labeled and Uncurated Unlabeled Data
abstract
Label-noise or curated unlabeled data are used to compensate for the assumption of clean labeled data in training the conditional generative adversarial network; however, satisfying such an extended assumption is occasionally laborious or impractical. As a step towards generative modeling accessible to everyone, we introduce a novel conditional image generation framework that accepts noisy-labeled and uncurated unlabeled data during training: (i) closed-set and open-set label noise in labeled data and (ii) closed-set and open-set unlabeled data. To combat it, we propose soft curriculum learning, which assigns instance-wise weights for adversarial training while assigning new labels for unlabeled data and correcting wrong labels for labeled data. Unlike popular curriculum learning, which uses a threshold to pick the training samples, our soft curriculum controls the effect of each training instance by using the weights predicted by the auxiliary classifier, resulting in the preservation of useful samples while ignoring harmful ones. Our experiments show that our approach outperforms existing semi-supervised and label-noise robust methods in terms of both quantitative and qualitative performance. In particular, the proposed approach matches the performance of (semi-)supervised GANs even with less than half the labeled data.1
Kai Katsumata, Duc Minh Vo, Tatsuya Harada, Hideki Nakayama
WACV4
2024 Revisiting Latent Space of GAN Inversion for Robust Real Image Editing
abstract
We present a generative adversarial network (GAN) inversion with high reconstruction and editing quality. GAN inversion algorithms with expressive latent spaces produce near-perfect inversion but are not robust to editing operations in a latent space, leading to undesirable edited images, a phenomenon known as the trade-off between reconstruction and editing quality. To cope with the trade-off, we revisit the hyperspherical prior of StyleGANs $\mathcal{Z}$ and propose to combine an extended space of $\mathcal{Z}$ with highly capable inversion algorithms. Our approach maintains the reconstruction quality of seminal GAN inversion methods while improving their editing quality owing to the constrained nature of $\mathcal{Z}$. Through comprehensive experiments with several GAN inversion algorithms, we demonstrate that our approach enhances the image editing quality in 2D/3D GANs.1
Kai Katsumata, Duc Minh Vo, Bei Liu 0001, Hideki Nakayama
WACV4
2024 Label Augmentation as Inter-class Data Augmentation for Conditional Image Synthesis with Imbalanced Data
abstract
Conditional image synthesis performs admirably when trained on well-constructed and balanced datasets. However, in practice, training datasets frequently contain minorities (i.e., a class with a few samples), known as imbalanced data, which causes difficulties in learning generative models. To address conditional image synthesis with imbalanced data, we analyze a diversity issue of label-preserving data augmentation and an affinity issue of non-label-preserving data augmentation. From this observation, we present label augmentation, which works as inter-class data augmentation that effectively augments data by predicting a new label for a given image using the prediction of a pretrained image classification model (i.e., probabilities for each class). We incorporate our label augmentation into the discriminator of a seminal conditional generative adversarial network (GAN) model, proposing Softlabel-GAN. Using class probabilities extracts class-invariant and shared features between similar classes, achieving data augmentation with high affinity and diversity. Our experiments on imbalanced datasets show that Softlabel-GAN produces images with high quality and diversity while being hardly affected by the number of samples in each class. Code: https://github.com/raven38/softlabel-gan.
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
WACV3
2023 A-CAP: Anticipation Captioning with Commonsense Knowledge
abstract
Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered set of images. To tackle this new task, we propose a model called A-CAP, which incorporates commonsense knowledge into a pre-trained vision-language model, allowing it to anticipate the caption. Through both qualitative and quantitative evaluations on a customized visual storytelling dataset, A-CAP out-performs other image captioning methods and establishes a strong baseline for anticipation captioning. We also address the challenges inherent in this task.
Duc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki Nakayama
CVPR4
2023 Partition-and-Debias: Agnostic Biases Mitigation via A Mixture of Biases-Specific Experts
abstract
Bias mitigation in image classification has been widely researched, and existing methods have yielded notable results. However, most of these methods implicitly assume that a given image contains only one type of known or unknown bias, failing to consider the complexities of real-world biases. We introduce a more challenging scenario, agnostic biases mitigation, aiming at bias removal regardless of whether the type of bias or the number of types is unknown in the datasets. To address this difficult task, we present the Partition-and-Debias (PnD) method that uses a mixture of biases-specific experts to implicitly divide the bias space into multiple subspaces and a gating module to find a consensus among experts to achieve debiased classification. Experiments on both public and constructed benchmarks demonstrated the efficacy of the PnD. Code is available at: https://github.com/Jiaxuan-Li/PnD.
Jiaxuan Li 0004, Duc Minh Vo, Hideki Nakayama
ICCV3
2023 Handwritten Text Generation with Character-Specific Encoding for Style Imitation
Jan Zdenek, Hideki Nakayama
ICDAR (2)2
2023 Indirect Adversarial Losses via an Intermediate Distribution for Training GANs
abstract
In this study, we consider the weak convergence characteristics of the Integral Probability Metrics (IPM) methods in training Generative Adversarial Networks (GANs). We first concentrate on a successful IPM-based GAN method that employs a repulsive version of the Maximum Mean Discrepancy (MMD) as the discriminator loss (called repulsive MMD-GAN). We reinterpret its repulsive metrics as an indirect discriminator loss function toward an intermediate distribution. This allows us to propose a novel generator loss via such an intermediate distribution based on our reinterpretation. Our indirect adversarial losses use a simple known distribution (i.e., the Normal or Uniform distribution in our experiments) to simulate indirect adversarial learning between three parts – real, fake, and intermediate distributions. Furthermore, we found the Kernelized Stein Discrepancy (KSD) from the IPM family as the adversarial loss function to avoid randomness from intermediate distribution samples because the target side (intermediate one) is sample-free in KSD. Experiments on several real-world datasets show that our methods can successfully train GANs with the intermediate-distribution-based KSD and MMD and can outperform previous loss metrics.
Duc Minh Vo, Hideki Nakayama
WACV3
2022 RNSum: A Large-Scale Dataset for Automatic Release Note Generation via Commit Logs Summarization
abstract
Hisashi Kamezawa, Noriki Nishida, Nobuyuki Shimizu, Takashi Miyazaki, Hideki Nakayama. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Hisashi Kamezawa, Noriki Nishida, Nobuyuki Shimizu, Takashi Miyazaki, Hideki Nakayama
ACL (1)5
2022 Weakly Supervised Formula Learner for Solving Mathematical Problems
abstract
Mathematical reasoning task is a subset of the natural language question answering task. Existing work suggested solving this task with a two-phase approach, where the model first predicts formulas from questions and then calculates answers from such formulas. This approach achieved desirable performance in existing work. However, its reliance on annotated formulas as intermediate labels throughout its training limited its application. In this work, we put forward the idea to enable models to learn optimal formulas autonomously. We proposed Weakly Supervised Formula Learner, a learning framework that drives the formula exploration with weak supervision from the final answers to mathematical problems. Our experiments are conducted on two representative mathematical reasoning datasets MathQA and Math23K. On MathQA, our method outperformed baselines trained on complete yet imperfect formula annotations. On Math23K, our method outperformed other weakly supervised learning methods.
Hideki Nakayama
COLING2
2022 OSSGAN: Open-Set Semi-Supervised Image Generation
abstract
We introduce a challenging training scheme of conditional GANs, called open-set semi-supervised image generation, where the training dataset consists of two parts: (i) labeled data and (ii) unlabeled data with samples belonging to one of the labeled data classes, namely, a closed-set, and samples not belonging to any of the labeled data classes, namely, an open-set. Unlike the existing semi-supervised image generation task, where unlabeled data only contain closed-set samples, our task is more general and lowers the data collection cost in practice by allowing open-set samples to appear. Thanks to entropy regularization, the classifier that is trained on labeled data is able to quantify sample-wise importance to the training of cGAN as confidence, allowing us to use all samples in un-labeled data. We design OSSGAN, which provides decision clues to the discriminator on the basis of whether an unlabeled image belongs to one or none of the classes of interest, smoothly integrating labeled and unlabeled data during training. The results of experiments on Tiny ImageNet and ImageNet show notable improvements over supervised Big-GAN and semi-supervised methods. Our code is available at https://github.com/raven38/OSSGAN.
Kai Katsumata, Duc Minh Vo, Hideki Nakayama
CVPR3
2022 NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge
abstract
Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowledge in the form of embeddings for any object's definition from Wiktionary, where we use in the retrieval image region features learned from a transformers model. We propose an end-to-end Novel Object Captioning with Retrieved vocabulary from External Knowledge method (NOC-REK), which simultaneously learns vocabulary retrieval and caption generation, successfully describing novel objects outside of the training dataset. Furthermore, our model eliminates the requirement for model retraining by simply updating the external knowledge whenever a novel object appears. Our comprehensive experiments on held-out COCO and Nocaps datasets show that our NOCREK is considerably effective against SOTAs.
Duc Minh Vo, Hong Chen 0017, Akihiro Sugimoto, Hideki Nakayama
CVPR4
2022 Character-centric Story Visualization via Visual Planning and Token Alignment
abstract
Story visualization advances the traditional text-to-image generation by enabling multiple image generation based on a complete story.This task requires machines to 1) understand long text inputs and 2) produce a globally consistent image sequence that illustrates the contents of the story.A key challenge of consistent story visualization is to preserve characters that are essential in stories.To tackle the challenge, we propose to adapt a recent work that augments Vector-Quantized Variational Autoencoders (VQ-VAE) with a text-tovisual-token (transformer) architecture.Specifically, we modify the text-to-visual-token module with a two-stage framework: 1) character token planning model that predicts the visual tokens for characters only; 2) visual token completion model that generates the remaining visual token sequence, which is sent to VQ-VAE for finalizing image generations.To encourage characters to appear in the images, we further train the two-stage framework with a character-token alignment objective.Extensive experiments and evaluations demonstrate that the proposed method excels at preserving characters and can produce higher quality image sequences compared with the strong baselines.Code can be found in https: //github.com/PlusLabNLP/VP-CSV
Hong Chen 0017, Rujun Han, Te-Lin Wu, Hideki Nakayama, Nanyun Peng 0001
EMNLP4
2022 StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning
abstract
Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tasks: Ranking, Rating and Reasoning.Given either a machine-generated or a human-written story, StoryER requires the machine to output 1) a preference score that corresponds to human preference, 2) specific ratings and their corresponding confidences and 3) comments for various aspects (e.g., opening, character-shaping).To support these tasks, we introduce a wellannotated dataset comprising (i) 100k ranked story pairs; and (ii) a set of 46k ratings and comments on various aspects of the story.We finetune Longformer-Encoder-Decoder (LED) on the collected dataset, with the encoder responsible for preference score and aspect prediction and the decoder for comment generation.Our comprehensive experiments result in a competitive benchmark for each task, showing the high correlation to human preference.In addition, we have witnessed the joint learning of the preference scores, the aspect ratings, and the comments brings gain in each single task.Our dataset and benchmarks are publicly available to advance the research of story evaluation tasks. 1
Hong Chen 0017, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao, Hideki Nakayama
EMNLP5
2022 Pixel to Binary Embedding Towards Robustness for CNNs
abstract
There are several problems with the robustness of Convolutional Neural Networks (CNNs). For example, the prediction of CNNs can be changed by adding a small magnitude of noise to an input, and the performances of CNNs are degraded when the distribution of input is shifted by a transformation never seen during training (e.g., the blur effect). There are approaches to replace pixel values with binary embeddings to tackle the problem of adversarial perturbations, which successfully improve robustness. In this work, we propose Pixel to Binary Embedding (P2BE) to improve the robustness of CNNs. P2BE is a learnable binary embedding method as opposed to previous hand-coded binary embedding methods. P2BE outperforms other binary embedding methods in robustness against adversarial perturbations and visual corruptions that are not shown during training.
Ikki Kishida, Hideki Nakayama
ICPR2
2022 DJMix: Unsupervised Task-agnostic Image Augmentation for Improving Robustness of Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) are vulnerable to unseen test-time noise on input images, such as defocus blur or JPEG-compression artifacts. Improving the robustness to such noise is important for real-world applications. In this paper, we propose DJMix, a training method for CNNs to obtain identical representations to a given image and its discretized one. As a result, CNNs trained with DJMix can ignore unnecessary details of inputs and become robust to input noise. We verify the effectiveness of our method on several datasets of various tasks, namely, classification, semantic segmentation, and object detection using clean and noisy test images.
Ryuichiro Hataya, Hideki Nakayama
IJCNN2
2022 Meta Approach to Data Augmentation Optimization
abstract
Data augmentation policies drastically improve the performance of image recognition tasks, especially when the policies are optimized for the target data and tasks. In this paper, we propose to optimize image recognition models and data augmentation policies simultaneously to improve the performance using gradient descent. Unlike prior methods, our approach avoids using proxy tasks or reducing search space, and can directly improve the validation performance. Our method achieves efficient and scalable training by approximating the gradient of policies by implicit gradient with Neumann series approximation. We demonstrate that our approach can improve the performance of various image classification tasks, including fine-grained image recognition, without using dataset-specific hyperparameter tuning.
Ryuichiro Hataya, Jan Zdenek, Kazuki Yoshizoe, Hideki Nakayama
WACV4
2022 PPCD-GAN: Progressive Pruning and Class-Aware Distillation for Large-Scale Conditional GANs Compression
abstract
We push forward neural network compression research by exploiting a novel challenging task of large-scale conditional generative adversarial networks (GANs) compression. To this end, we propose a gradually shrinking GAN (PPCD-GAN) by introducing progressive pruning residual block (PP-Res) and class-aware distillation. The PP-Res is an extension of the conventional residual block where each convolutional layer is followed by a learnable mask layer to progressively prune network parameters as training proceeds. The class-aware distillation, on the other hand, enhances the stability of training by transferring immense knowledge from a well-trained teacher model through instructive attention maps. We train the pruning and distillation processes simultaneously on a well-known GAN architecture in an end-to-end manner. After training, all redundant parameters as well as the mask layers are discarded, yielding a lighter network while retaining the performance. We comprehensively illustrate, on ImageNet 128 × 128 dataset, PPCD-GAN reduces up to 5.2 ×(81%) parameters against state-of-the-arts while keeping better performance.
Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
WACV3
2021 Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling
abstract
Visual storytelling is a task of generating relevant and interesting stories for given image sequences. In this work we aim at increasing the diversity of the generated stories while preserving the informative content from the images. We propose to foster the diversity and informativeness of a generated story by using a concept selection module that suggests a set of concept candidates. Then, we utilize a large scale pre-trained model to convert concepts and images into full stories. To enrich the candidate concepts, a commonsense knowledge graph is created for each image sequence from which the concept candidates are proposed. To obtain appropriate concepts from the graph, we propose two novel modules that consider the correlation among candidate concepts and the image-concept correlation. Extensive automatic and human evaluation results demonstrate that our model can produce reasonable concepts. This enables our model to outperform the previous models by a large margin on the diversity and informativeness of the story, while retaining the relevance of the story to the image sequence.
Hong Chen 0017, Yifei Huang 0002, Hiroya Takamura, Hideki Nakayama
AAAI4
2021 Visualizing Association in Exemplar-Based Classification
abstract
Recent progress in deep learning has enhanced image classification performance. However, classification using deep convolutional neural networks lacks interpretability. To solve this problem, we propose a novel method of explainable classification; this method uses images representing each image class, which we call exemplars. Our method comprises encoder-decoder models (association networks) and a classifier. First, the association networks transform each input image into an image that a deep neural network associates, which we call an associative image. Then, the image-level similarity between the associative images and the exemplars is used as a feature for classification. This similarity explains the decision of the classifiers. We conducted experiments using CIFAR-10, CIFAR-100, and STL-10 and demonstrated our classifier’s interpretability through the proposed visualization technique.
Taiga Kashima, Ryuichiro Hataya, Hideki Nakayama
ICASSP3
2021 Semantic Image Synthesis from Inaccurate and Coarse Masks
abstract
Semantic image synthesis is an image-to-image translation problem where the goal is to learn mapping from semantic segmentation masks to corresponding photorealistic images. However, conventional semantic image synthesis methods require numerous pairs of correct semantic masks and real images, and collecting these pairs is not always possible. To address this issue, we propose a smoothing method, which we call local label smoothing (LLS), that incorporates label smoothing per small patch of an input mask to learn mapping from masks to images even when semantic masks are inaccurate. Furthermore, we also propose an extended method for coarse masks. We demonstrate the advantage of the proposed methods over existing methods to deal with noisy masks on several datasets.
Kai Katsumata, Hideki Nakayama
ICASSP2
2021 Open-Set Domain Generalization VIA Metric Learning
abstract
In this study, we address open-set domain generalization, which aims to reject unknown class samples while classifying known class samples in unseen domains. Conventional domain generalization has the problem of unknown class samples being classified as known classes because domain generalization methods align feature distributions without distinction between known and unknown classes. To tackle this problem, we propose a decoupling loss that diffuses the feature representations of unknown samples. The loss allows us to construct a feature space that can better distinguish unknown samples. We demonstrate the effectiveness of decoupling loss using open-set domain generalization benchmarks.
Kai Katsumata, Ikki Kishida, Ayako Amma, Hideki Nakayama
ICIP4
2021 DCT-based Fast Spectral Convolution for Deep Convolutional Neural Networks
abstract
Spectral representations have been introduced into deep convolutional neural networks (CNNs) mainly for accelerating convolutions and mitigating information loss. However, repeated domain transformations and complex arithmetic of commonly-used Fourier transform (DFT, FFT) seriously limit the applicability of spectral networks. In contrast, discrete cosine transform (DCT)-based methods are more promising owing to computations with merely real numbers. Hence in this work, we investigate the convolution theorem of DCT and propose a faster spectral convolution method for CNNs. First, we transform the input feature map and convolutional kernel into the frequency domain via DCT. We then perform element-wise multiplication between the spectral feature map and kernel, which is mathematically equivalent to symmetric convolution in the spatial domain but much cheaper than the straightforward spatial convolution. Since DCT only involves real arithmetic, the computational complexity of our method is significantly smaller than the traditional FFT-based spectral convolution. Besides, we introduce a network optimization strategy to suppress repeated domain transformations leveraging the intrinsically extended kernels. Furthermore, we present a partial symmetry breaking strategy with spectral dropout to mitigate the performance degradation caused by kernel symmetry. Experimental results demonstrate that compared with traditional spatial and spectral methods, our proposed DCT-based spectral convolution effectively accelerates the networks while achieving comparable accuracy.
Hideki Nakayama
IJCNN2
2021 GraphPlan: Story Generation by Planning with Event Graph
abstract
Story generation is a task that aims to automatically generate a meaningful story.This task is challenging because it requires high-level understanding of the semantic meaning of sentences and causality of story events.Naive sequence-to-sequence models generally fail to acquire such knowledge, as it is difficult to guarantee logical correctness in a text generation model without strategic planning.In this study, we focus on planning a sequence of events assisted by event graphs and use the events to guide the generator.Rather than using a sequence-to-sequence model to output a sequence, as in some existing works, we propose to generate an event sequence by walking on an event graph.The event graphs are built automatically based on the corpus.To evaluate the proposed approach, we incorporate human participation, both in event planning and story generation.Based on the largescale human annotation results, our proposed approach has been shown to provide more logically correct event sequences and stories compared with previous approaches.
Hong Chen 0017, Raphael Shu, Hiroya Takamura, Hideki Nakayama
INLG4
2021 JokerGAN: Memory-Efficient Model for Handwritten Text Generation with Text Line Awareness
abstract
Collecting labeled data for training of models for image recognition problems, including handwritten text recognition (HTR), is a tedious and expensive task. Recent work on handwritten text generation shows that generative models can be used as a data augmentation method to improve the performance of HTR systems.
Jan Zdenek, Hideki Nakayama
ACM Multimedia2
2021 Object Recognition with Continual Open Set Domain Adaptation for Home Robot
abstract
Object recognition ability is indispensable for robots to act like humans in a home environment. For example, when considering an object searching task, humans can recognize a naturally arranged object previously held in their hands while ignoring never observed objects. Even in such a simple task, we need to deal with three complex problems: domain adaptation, open-set recognition, and continual learning. However, most existing datasets are simplified to focus on one problem and do not measure the object recognition ability for home robots when multiple problems are simultaneously present. In this paper, we propose the COSDA-HR (Continual Open Set Domain Adaptation for Home Robot) dataset that requires dealing with the above three problems simultaneously. The COSDA-HR dataset focuses particularly on the scenario in which naturally arranged objects in a room are recognized by training with handheld objects towards the goal of creating a user-friendly teaching system for home robots. We provide various baselines to address the problems in the COSDA-HR dataset by combining state-of-the-art methods from each research area and analyze the limitations of such simple combinations. We consider that it is necessary to study the methods of handling multiple problems simultaneously instead of solving each problem to realize practical object recognition systems for home robots.
Ikki Kishida, Hong Chen 0017, Masaki Baba, Jiren Jin, Ayako Amma, Hideki Nakayama
WACV6
2021 MADGAN: unsupervised medical anomaly detection GAN using multiple adjacent brain MRI slice reconstruction
abstract
BACKGROUND: Unsupervised learning can discover various unseen abnormalities, relying on large-scale unannotated medical images of healthy subjects. Towards this, unsupervised methods reconstruct a 2D/3D single medical image to detect outliers either in the learned feature space or from high reconstruction loss. However, without considering continuity between multiple adjacent slices, they cannot directly discriminate diseases composed of the accumulation of subtle anatomical anomalies, such as Alzheimer's disease (AD). Moreover, no study has shown how unsupervised anomaly detection is associated with either disease stages, various (i.e., more than two types of) diseases, or multi-sequence magnetic resonance imaging (MRI) scans. RESULTS: We propose unsupervised medical anomaly detection generative adversarial network (MADGAN), a novel two-step method using GAN-based multiple adjacent brain MRI slice reconstruction to detect brain anomalies at different stages on multi-sequence structural MRI: (Reconstruction) Wasserstein loss with Gradient Penalty + 100 [Formula: see text] loss-trained on 3 healthy brain axial MRI slices to reconstruct the next 3 ones-reconstructs unseen healthy/abnormal scans; (Diagnosis) Average [Formula: see text] loss per scan discriminates them, comparing the ground truth/reconstructed slices. For training, we use two different datasets composed of 1133 healthy T1-weighted (T1) and 135 healthy contrast-enhanced T1 (T1c) brain MRI scans for detecting AD and brain metastases/various diseases, respectively. Our self-attention MADGAN can detect AD on T1 scans at a very early stage, mild cognitive impairment (MCI), with area under the curve (AUC) 0.727, and AD at a late stage with AUC 0.894, while detecting brain metastases on T1c scans with AUC 0.921. CONCLUSIONS: Similar to physicians' way of performing a diagnosis, using massive healthy training data, our first multiple MRI slice reconstruction approach, MADGAN, can reliably predict the next 3 slices from the previous 3 ones only for unseen healthy images. As the first unsupervised various disease diagnosis, MADGAN can reliably detect the accumulation of subtle anatomical anomalies and hyper-intense enhancing lesions, such as (especially late-stage) AD and brain metastases on multi-sequence MRI scans.
Leonardo Rundo, Kohei Murao, Tomoyuki Noguchi, Yuki Shimahara, Zoltán Ádám Milacski, Saori Koshino, Evis Sala, Hideki Nakayama, Shin'ichi Satoh 0001
BMC Bioinform.9
2020 Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta Posterior
abstract
Although neural machine translation models reached high translation quality, the autoregressive nature makes inference difficult to parallelize and leads to high translation latency. Inspired by recent refinement-based approaches, we propose LaNMT, a latent-variable non-autoregressive model with continuous latent variables and deterministic inference procedure. In contrast to existing approaches, we use a deterministic inference algorithm to find the target sequence that maximizes the lowerbound to the log-probability. During inference, the length of translation automatically adapts itself. Our experiments show that the lowerbound can be greatly increased by running the inference algorithm, resulting in significantly improved translation quality. Our proposed model closes the performance gap between non-autoregressive and autoregressive approaches on ASPEC Ja-En dataset with 8.6x faster decoding. On WMT'14 En-De dataset, our model narrows the gap with autoregressive baseline to 2.0 BLEU points with 12.5x speedup. By decoding multiple initial latent variables in parallel and rescore using a teacher model, the proposed model further brings the gap down to 1.0 BLEU point on WMT'14 En-De task with 6.8x speedup.
Raphael Shu, Jason Lee 0002, Hideki Nakayama, Kyunghyun Cho
AAAI3
2020 Graph-Based Heuristic Search for Module Selection Procedure in Neural Module Network
Hideki Nakayama
ACCV (3)2
2020 Single Model Ensemble using Pseudo-Tags and Distinct Vectors
abstract
Model ensemble techniques often increase task performance in neural networks; however, they require increased time, memory, and management effort.In this study, we propose a novel method that replicates the effects of a model ensemble with a single model.Our approach creates K-virtual models within a single parameter space using K-distinct pseudotags and K-distinct vectors.Experiments on text classification and sequence labeling tasks on several datasets demonstrate that our method emulates or outperforms a traditional model ensemble with 1/K-times fewer parameters.
Ryosuke Kuwabara, Jun Suzuki 0001, Hideki Nakayama
ACL3
2020 Supervised Visual Attention for Multimodal Neural Machine Translation
abstract
This paper proposed a supervised visual attention mechanism for multimodal neural machine translation (MNMT), trained with constraints based on manual alignments between words in a sentence and their corresponding regions of an image.The proposed visual attention mechanism captures the relationship between a word and an image region more precisely than a conventional visual attention mechanism trained through MNMT in an unsupervised manner.Our experiments on English-German and German-English translation tasks using the Multi30k dataset and on English-Japanese and Japanese-English translation tasks using the Flickr30k Entities JP dataset show that a Transformer-based MNMT model can be improved by incorporating our proposed supervised visual attention mechanism and that further improvements can be achieved by combining it with a supervised cross-lingual attention mechanism (up to +1.61 BLEU, +1.
Tetsuro Nishihara, Akihiro Tamura, Takashi Ninomiya, Yutaro Omote, Hideki Nakayama
COLING5
2020 Faster AutoAugment: Learning Augmentation Strategies Using Backpropagation
Ryuichiro Hataya, Jan Zdenek, Kazuki Yoshizoe, Hideki Nakayama
ECCV (25)4
2020 A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses
abstract
In real-world dialogue, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions.Non-verbal responses also play an important role in social interactions.In this paper, we propose a visuallygrounded first-person dialogue (VFD) dataset with verbal and non-verbal responses.The VFD dataset provides manually annotated (1) first-person images of agents, (2) utterances of human speakers, (3) eye-gaze locations of the speakers, and (4) the agents' verbal and nonverbal responses.We present experimental results obtained using the proposed VFD dataset and recent neural network models (e.g., BERT, ResNet).The results demonstrate that firstperson vision helps neural network models correctly understand human intentions, and the production of non-verbal responses is a challenging task like that of verbal responses.Our dataset is publicly available 1 .1 https://randd.yahoo.co.jp/en/softwaredataU: これのLはないのかしら V: 同じ服がたくさんあるからどれかはLじゃないかな N: 同じ服のサイズをチェックする -------------------------U: I wonder if there is an L for this.V: We have a lot of the same clothes, so I'm guessing one of them is an L
Hisashi Kamezawa, Noriki Nishida, Nobuyuki Shimizu, Takashi Miyazaki, Hideki Nakayama
EMNLP (1)5
2020 Unsupervised Visual Relationship Inference
abstract
Visual relationship inference is an essential research area for image understanding. Owing to the recent advancement of deep learning, significant signs of progress have been made in this challenging area. Standard approaches attempt to recognize visual relationships based on supervised learning by employing a carefully annotated dataset, in which images, triplets (subject-predicate-object), and bounding boxes are attached. However, preparing a large-scale dataset is very time consuming. This study proposes a novel method to infer visual relationships without image-triplet pairs. Our method tries to keep cycle consistency and plausibility of the inferred triplets. Our experimental results demonstrate that this method can infer predicates between objects in unpaired settings, and also achieving promising results using triplets parsed from external image descriptions.
Taiga Kashima, Kento Masui, Hideki Nakayama
ICIP3
2020 A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking
abstract
Visually-grounded natural language processing has become an important research direction in the past few years. However, majorities of the available cross-modal resources (e.g., image-caption datasets) are built in English and cannot be directly utilized in multilingual or non-English scenarios. In this study, we present a novel multilingual multimodal corpus by extending the Flickr30k Entities image-caption dataset with Japanese translations, which we name Flickr30k Entities JP (F30kEnt-JP). To the best of our knowledge, this is the first multilingual image-caption dataset where the captions in the two languages are parallel and have the shared annotations of many-to-many phrase-to-region linking. We believe that phrase-to-region as well as phrase-to-phrase supervision can play a vital role in fine-grained grounding of language and vision, and will promote many tasks such as multilingual image captioning and multimodal machine translation. To verify our dataset, we performed phrase localization experiments in both languages and investigated the effectiveness of our Japanese annotations as well as multilingual learning realized by our dataset.
Hideki Nakayama, Akihiro Tamura, Takashi Ninomiya
LREC1
2020 Efficient Base Class Selection Algorithms for Few-Shot Classification
abstract
Few-shot classification is a task to learn a classifier for novel classes with a limited number of examples on top of the known base classes which have a sufficient number of examples. In recent years, significant progress has been achieved on this task. However, despite the importance of selecting the base classes themselves for better knowledge transfer, few works have paid attention to this point. In this paper, we propose two types of base class selection algorithms that are suitable for few-shot classification tasks. One is based on the thesaurus-tree structure of class names, and the other is based on word embeddings. In our experiments using representative few-shot learning methods on the ILSVRC dataset, we show that these two algorithms can significantly improve the performance compared to a naive class selection method. Moreover, they do not require high computational and memory costs, which is an important advantage to scale to a very large number of base classes.
Takumi Ohkuma, Hideki Nakayama
ICMR2
2020 Erasing Scene Text with Weak Supervision
abstract
Scene text erasing is a task of removing text from natural scene images, which has been gaining attention in recent years. The main motivation is to conceal private information such as license plate numbers, and house nameplates that can appear in images. In this work, we propose a method for scene text erasing that approaches the problem as a general inpainting task. In contrast to previous methods, which require pairs of original images containing text and images from which the text has been removed, our method does not need corresponding image pairs for training. We use a separately trained scene text detector and an inpainting network. The scene text detector predicts segmentation maps of text instances which are then used as masks for the inpainting network. The network for inpainting, trained on a large-scale image dataset, fills in masked out regions in an input image and generates a final image in which the original text is no longer present. The results show that our method is able to successfully remove text and fill in the created holes to produce natural-looking images.
Jan Zdenek, Hideki Nakayama
WACV2
2020 Unsupervised Discourse Constituency Parsing Using Viterbi EM
abstract
In this paper, we introduce an unsupervised discourse constituency parsing algorithm. We use Viterbi EM with a margin-based criterion to train a span-based discourse parser in an unsupervised manner. We also propose initialization methods for Viterbi training of discourse constituents based on our prior knowledge of text structures. Experimental results demonstrate that our unsupervised parser achieves comparable or even superior performance to fully supervised parsers. We also investigate discourse constituents that are learned by our method.
Noriki Nishida, Hideki Nakayama
Trans. Assoc. Comput. Linguistics2
2019 Synthesizing Diverse Lung Nodules Wherever Massively: 3D Multi-Conditional GAN-Based CT Image Augmentation for Object Detection
abstract
Accurate Computer-Assisted Diagnosis, relying on large-scale annotated pathological images, can alleviate the risk of overlooking the diagnosis. Unfortunately, in medical imaging, most available datasets are small/fragmented. To tackle this, as a Data Augmentation (DA) method, 3D conditional Generative Adversarial Networks (GANs) can synthesize desired realistic/diverse 3D images as additional training data. However, no 3D conditional GAN-based DA approach exists for general bounding box-based 3D object detection, while it can locate disease areas with physicians' minimum annotation cost, unlike rigorous 3D segmentation. Moreover, since lesions vary in position/size/attenuation, further GAN-based DA performance requires multiple conditions. Therefore, we propose 3D Multi-Conditional GAN (MCGAN) to generate realistic/diverse 32 × 32 × 32 nodules placed naturally on lung Computed Tomography images to boost sensitivity in 3D object detection. Our MCGAN adopts two discriminators for conditioning: the context discriminator learns to classify real vs synthetic nodule/surrounding pairs with noise box-centered surroundings; the nodule discriminator attempts to classify real vs synthetic nodules with size/attenuation conditions. The results show that 3D Convolutional Neural Network-based detection can achieve higher sensitivity under any nodule size/attenuation at fixed False Positive rates and overcome the medical data paucity with the MCGAN-generated realistic nodules-even expert physicians fail to distinguish them from the real ones in Visual Turing Test.
Yoshiro Kitamura, Akira Kudo, Akimichi Ichinose, Leonardo Rundo, Yujiro Furukawa, Kazuki Umemoto, Yuanzhong Li, Hideki Nakayama
3DV9
2019 Generating Diverse Translations with Sentence Codes
abstract
Users of machine translation systems may desire to obtain multiple candidates translated in different ways.In this work, we attempt to obtain diverse translations by using sentence codes to condition the sentence generation.We describe two methods to extract the codes, either with or without the help of syntax information.For diverse generation, we sample multiple candidates, each of which conditioned on a unique code.Experiments show that the sampled translations have much higher diversity scores when using reasonable sentence codes, where the translation quality is still on par with the baselines even under strong constraint imposed by the codes.In qualitative analysis, we show that our method is able to generate paraphrase translations with drastically different structures.The proposed approach can be easily adopted to existing translation systems as no modification to the model is required.
Raphael Shu, Hideki Nakayama, Kyunghyun Cho
ACL (1)2
2019 Learning More with Less: Conditional PGGAN-based Data Augmentation for Brain Metastases Detection Using Highly-Rough Annotation on MR Images
abstract
Accurate Computer-Assisted Diagnosis, associated with proper data wrangling, can alleviate the risk of overlooking the diagnosis in a clinical environment. Towards this, as a Data Augmentation (DA) technique, Generative Adversarial Networks (GANs) can synthesize additional training data to handle the small/fragmented medical imaging datasets collected from various scanners; those images are realistic but completely different from the original ones, filling the data lack in the real image distribution. However, we cannot easily use them to locate disease areas, considering expert physicians' expensive annotation cost. Therefore, this paper proposes Conditional Progressive Growing of GANs (CPGGANs), incorporating highly-rough bounding box conditions incrementally into PGGANs to place brain metastases at desired positions/sizes on 256 X 256 Magnetic Resonance (MR) images, for Convolutional Neural Network-based tumor detection; this first GAN-based medical DA using automatic bounding box annotation improves the training robustness. The results show that CPGGAN-based DA can boost 10% sensitivity in diagnosis with clinically acceptable additional False Positives. Surprisingly, further tumor realism, achieved with additional normal brain MR images for CPGGAN training, does not contribute to detection performance, while even three physicians cannot accurately distinguish them from the real ones in Visual Turing Test.
Kohei Murao, Tomoyuki Noguchi, Yusuke Kawata, Fumiya Uchiyama, Leonardo Rundo, Hideki Nakayama, Shin'ichi Satoh 0001
CIKM7
2019 LOL: Learning To Optimize Loss Switching Under Label Noise
abstract
Deep convolutional neural networks excel in image recognition, but they are also known to be fragile to label corruption. To mitigate this problem, we propose to dynamically switch two loss functions, categorical cross entropy and mean absolute error, to exploit their complementary advantages. We employ the bilevel programming approach to simultaneously optimize base CNNs and the weights of two loss functions. Our proposed method only requires little modification in the optimization process of the original supervised problem and is applicable to a wide variety of networks under label corruption. Further, our approach achieves on-par results with other state-of-the-art noise-tolerant learning methods.
Ryuichiro Hataya, Hideki Nakayama
ICIP2
2019 DCT Based Information-Preserving Pooling for Deep Neural Networks
abstract
Pooling is used in most of deep convolutional neural networks as a feature downsampling method, in order to reduce computation complexity and increase the receptive field size. Since the traditional max/average pooling layers tend to cause severe information loss, spectral pooling with discrete Fourier transform (DFT) is considered as an advisable alternative. In this paper, we propose a novel 2D-discrete cosine transform (2D-DCT) based pooling method for deep neural networks. Due to the energy compaction property, DCT pooling preserves considerably more information than DFT. Moreover, inspired by the separability of DCT, we precompute the transform matrices and embed them into linear layers for parallelization and acceleration on GPUs. Experimental results indicate that the proposed DCT based pooling layer outperforms previous pooling methods on multiple image classification datasets with negligible extra time consumption.
Hideki Nakayama
ICIP2
2019 Bipolar Gan: Double Check the Solution Space and Lighten False Positive Errors in Generative Adversarial Nets
abstract
Generative Adversarial Nets (GAN) and its variations gain their popularity both in application scenarios and the research front. In this paper, we proposed a novel approach which is compatible with previous methods. It can improve the quality of generated images using both Authenticity Discriminator and Falsity Discriminator to double check the solution space. Our experiments exhibited the feasibility of our solution and expressed its ability to improve recently proposed methods.
Rui Yang 0021, Hideki Nakayama
ICIP2
2019 Empirical Study of Easy and Hard Examples in CNN Training
Ikki Kishida, Hideki Nakayama
ICONIP (4)2
2019 Shifted Spatial-Spectral Convolution for Deep Neural Networks
abstract
Deep convolutional neural networks (CNNs) extract local features and learn spatial representations via convolutions in the spatial domain. Beyond the spatial information, some works also manage to capture the spectral information in the frequency domain by domain switching methods like discrete Fourier transform (DFT) and discrete cosine transform (DCT). However, most works only pay attention to a single domain, which is prone to ignoring other important features. In this work, we propose a novel network structure to combine spatial and spectral convolutions, and extract features in both spatial and frequency domains. The input channels are divided into two groups for spatial and spectral representations respectively, and then integrated for feature fusion. Meanwhile, we design a channel-shifting mechanism to ensure both spatial and spectral information of every channel are equally and adequately obtained throughout the deep networks. Experimental results demonstrate that compared with state-of-the-art CNN models in a single domain, our shifted spatial-spectral convolution based networks achieve better performance on image classification datasets including CIFAR10, CIFAR100 and SVHN, with considerably fewer parameters.
Hideki Nakayama
MMAsia2
2019 USE-Net: Incorporating Squeeze-and-Excitation blocks into U-Net for prostate zonal segmentation of multi-institutional MRI datasets
Leonardo Rundo, Yudai Nagano, Ryuichiro Hataya, Carmelo Militello, Andrea Tangherloni, Marco S. Nobile, Claudio Ferretti, Daniela Besozzi, Maria Carla Gilardi, Salvatore Vitabile, Giancarlo Mauri, Hideki Nakayama, Paolo Cazzaniga
Neurocomputing14
2018 Semantic Aware Attention Based Deep Object Co-segmentation
Hong Chen 0017, Yifei Huang 0002, Hideki Nakayama
ACCV (4)3
2018 Compressing Word Embeddings via Deep Compositional Code Learning
Raphael Shu, Hideki Nakayama
ICLR (Poster)2
2018 Incorporating Semantic Attention in Video Description Generation
Natsuda Laokulrat, Naoaki Okazaki, Hideki Nakayama
LREC3
2018 Augmenting Image Question Answering Dataset by Exploiting Image Captions
Masashi Yokota, Hideki Nakayama
LREC2
2018 PoB: Toward Reasoning Patterns of Beauty in Image Data
abstract
Aiming to develop of computational grammar system for visual information, we design a 4-tier framework that consists of four levels of 'visual grammar of images.' As a first step of realization, we propose a new dataset, named the PoB dataset, in which each image is annotated with multiple labels of armature patterns that compose the pictorial scene. The PoB dataset includes of a 10,000-painting dataset for art and a 4,959-image dataset for photography. In this paper, we discuss the consistency analysis of our dataset and its applicability. We also demonstrate how the armature patterns in the PoB dataset are useful in assessing aesthetic quality of images, and how well a deep learning algorithm can recognize these patterns. This paper seeks to set a new direction in image understanding with a more holistic approach beyond discrete objects and in aesthetic reasoning with a more interpretative way.
Diep Thi Ngoc Nguyen, Hideki Nakayama, Naoaki Okazaki, Tatsuya Sakaeda
ACM Multimedia2
2018 Deep Learning for Forecasting Stock Returns in the Cross-Section
Masaya Abe, Hideki Nakayama
PAKDD (1)2
2018 Coherence Modeling Improves Implicit Discourse Relation Recognition
abstract
The research described in this paper examines how to learn linguistic knowledge associated with discourse relations from unlabeled corpora.We introduce an unsupervised learning method on text coherence that could produce numerical representations that improve implicit discourse relation recognition in a semi-supervised manner.We also empirically examine two variants of coherence modeling: orderoriented and topic-oriented negative sampling, showing that, of the two, topicoriented negative sampling tends to be more effective.
Noriki Nishida, Hideki Nakayama
SIGDIAL Conference2
2017 Bag of Local Convolutional Triplets for Script Identification in Scene Text
abstract
The increasing interest in scene text reading in multilingual environments raises the need to recognize and distinguish between different writing systems. In this paper, we propose a novel method for script identification in scene text using triplets of local convolutional features in combination with the traditional bag-of-visual-words model. Feature triplets are created by making combinations of descriptors extracted from local patches of the input images using a convolutional neural network. This approach allows us to generate a more descriptive codeword dictionary for the bag-of-visual-words model, as the low discriminative power of weak descriptors is enhanced by other descriptors in a triplet. The proposed method is evaluated on two public benchmark datasets for scene text script identification and a public dataset for script identification in video captions. The experiments demonstrate that our method outperforms the baseline and yields competitive results on all three datasets.
Jan Zdenek, Hideki Nakayama
ICDAR2
2017 Word Ordering as Unsupervised Learning Towards Syntactically Plausible Word Representations
abstract
The research question we explore in this study is how to obtain syntactically plausible word representations without using human annotations. Our underlying hypothesis is that word ordering tests, or linearizations, is suitable for learning syntactic knowledge about words. To verify this hypothesis, we develop a differentiable model called Word Ordering Network (WON) that explicitly learns to recover correct word order while implicitly acquiring word embeddings representing syntactic knowledge. We evaluate the word embeddings produced by the proposed method on downstream syntax-related tasks such as part-of-speech tagging and dependency parsing. The experimental results demonstrate that the WON consistently outperforms both order-insensitive and order-sensitive baselines on these tasks.
Noriki Nishida, Hideki Nakayama
IJCNLP(1)2
2017 Recurrent Visual Relationship Recognition with Triplet Unit
abstract
The task of visual relationship recognition (VRR) is recognizing multiple objects and their relationships in an image. A fundamental difficulty of this task is class-number scalability, since the number of possible relationships we need to consider causes combinatorial explosion. Another difficulty of this task is modeling how to avoid outputting semantically redundant relationships. To overcome these challenges, this paper proposes a novel architecture with a recurrent neural network (RNN) and triplet unit (TU). The RNN allows our model to be optimized for outputting a sequence of relationships. By optimizing our model to a semantically diverse relationship sequence, we increase the variety in output relationships. At each step of the RNN, our TU enables the model to classify a relationship while achieving class-number scalability by decomposing a relationship into a subject-predicate-object (SPO) triplet. We evaluate our model on various datasets and compare the results to a baseline. These experimental results show our model's superior recall and precision with fewer predictions compared to the baseline, even as it produces greater variety in relationships.
Kento Masui, Akiyoshi Ochiai, Shintaro Yoshizawa, Hideki Nakayama
ISM4
2017 Zero-resource machine translation by multimodal encoder-decoder network with multimedia pivot
abstract
We propose an approach to build a neural machine translation system with no supervised resources (i.e., no parallel corpora) using multimodal embedded representation over texts and images. Based on the assumption that text documents are often likely to be described with other multimedia information (e.g., images) somewhat related to the content, we try to indirectly estimate the relevance between two languages. Using multimedia as the “pivot”, we project all modalities into one common hidden space where samples belonging to similar semantic concepts should come close to each other, whatever the observed space of each sample is. This modality-agnostic representation is the key to bridging the gap between different modalities. Putting a decoder on top of it, our network can flexibly draw the outputs from any input modality. Notably, in the testing phase, we need only source language texts as the input for translation. In experiments, we tested our method on two benchmarks to show that it can achieve reasonable translation performance. We compared and investigated several possible implementations and found that an end-to-end model that simultaneously optimized both rank loss in multimodal encoders and cross-entropy loss in decoders performed the best.
Hideki Nakayama, Noriki Nishida
Mach. Transl.1
2016 Generating Video Description using Sequence-to-sequence Model with Temporal Attention
abstract
Automatic video description generation has recently been getting attention after rapid advancement in image caption generation. Automatically generating description for a video is more challenging than for an image due to its temporal dynamics of frames. Most of the work relied on Recurrent Neural Network (RNN) and recently attentional mechanisms have also been applied to make the model learn to focus on some frames of the video while generating each word in a describing sentence. In this paper, we focus on a sequence-to-sequence approach with temporal attention mechanism. We analyze and compare the results from different attention model configuration. By applying the temporal attention mechanism to the system, we can achieve a METEOR score of 0.310 on Microsoft Video Description dataset, which outperformed the state-of-the-art system so far.
Natsuda Laokulrat, Sang Phan Le, Noriki Nishida, Raphael Shu, Yo Ehara, Naoaki Okazaki, Yusuke Miyao, Hideki Nakayama
COLING8
2016 Annotation order matters: Recurrent Image Annotator for arbitrary length image tagging
abstract
Automatic image annotation has been an important research topic in facilitating large scale image management and retrieval. Existing methods focus on learning image-tag correlation or correlation between tags to improve annotation accuracy. However, most of these methods evaluate their performance using top-k retrieval performance, where k is fixed. Although such setting gives convenience for comparing different methods, it is not the natural way that humans annotate images. The number of annotated tags should depend on image contents. Inspired by the recent progress in machine translation and image captioning, we propose a novel Recurrent Image Annotator (RIA) model that forms image annotation task as a sequence generation problem so that RIA can natively predict the proper length of tags according to image contents. We evaluate the proposed model on various image annotation datasets. In addition to comparing our model with existing methods using the conventional top-k evaluation measures, we also provide our model as a high quality baseline for the arbitrary length image tagging task. Moreover, the results of our experiments show that the order of tags in training phase has a great impact on the final annotation performance.
Jiren Jin, Hideki Nakayama
ICPR2
2015 Image-Mediated Learning for Zero-Shot Cross-Lingual Document Retrieval
abstract
We propose an image-mediated learning approach for cross-lingual document retrieval where no or only a few parallel corpora are available.Using the images in image-text documents of each language as the hub, we derive a common semantic subspace bridging two languages by means of generalized canonical correlation analysis.For the purpose of evaluation, we create and release a new document dataset consisting of three types of data (English text, Japanese text, and images).Our approach substantially enhances retrieval accuracy in zero-shot and few-shot scenarios where text-to-text examples are scarce.
Ruka Funaki, Hideki Nakayama
EMNLP2
2015 Unsupervised Cosegmentation based on Global Graph Matching
abstract
Cosegmentation is defined as the task of segmenting a common object from multiple images. Hitherto, graph matching has been known as a promising approach because of its flexibility in matching deformable objects and regions, and several methods based on this approach have been proposed. However, candidate foregrounds obtained by a local matching algorithm in previous methods tend to include false-positive areas, particularly when visually similar backgrounds (e.g., sky) commonly appear across images.
Takanori Tamanaha, Hideki Nakayama
ACM Multimedia2
2015 Multimodal Gesture Recognition Using Multi-stream Recurrent Neural Network
Noriki Nishida, Hideki Nakayama
PSIVT2
2014 Unsupervised Visual Domain Adaptation Using Auxiliary Information in Target Domain
abstract
We propose a novel approach for unsupervised visual domain adaptation that exploits auxiliary information in a target domain. The key idea is to embed data in the target domain into a subspace where samples are better organized, expecting auxiliary information to serve as a somewhat semantically related signal. Specifically, we apply partial least squares (PLS) to RGB image features and corresponding depth features captured at the same time. Thus, we can improve the performance of domain adaptation without any help from manual annotation in the target domain. In experiments, we tested our approach with two state-of-the-art subspace based domain adaptation methods and show that, our method consistently improves the classification accuracy.
Masaya Okamoto, Hideki Nakayama
ISM2
2013 Efficient Discriminative Convolution Using Fisher Weight Map
Hideki Nakayama
BMVC1
2013 Augmenting descriptors for fine-grained visual categorization using polynomial embedding
abstract
Fine-grained visual categorization (FGVC), which is a relatively new research area, distinguishes conceptually and visually similar categories such as plant and animal species. While FGVC is expected to lead to many task-specific practical applications, it is known as an extremely difficult problem because interclass variations are often quite subtle. We believe that the key to FGVC is improving local descriptors to enhance discriminative power at the local patch-level. While the pooling strategy of descriptors has been intensively improved for bag-of-visual-words (BoVW) based image representations, the descriptors themselves are often untouched. In this paper, we propose a descriptor augmentation method that utilizes polynomial embedding and supervised dimensionality reduction. Since our method provides moderate-sized compressed descriptors, it can be naturally integrated with off-the-shelf BoVW techniques. In experiments, we show that our method achieves state-of-the-art performance on standard FGVC datasets, Caltech-Birds, and Oxford-Flowers.
Hideki Nakayama
ICME1
2010 Evaluation of dimensionality reduction methods for image auto-annotation
abstract
Image auto-annotation is a challenging task in computer vision. The goal of this task is to predict multiple words for generic images automatically. Recent state-of-theart methods are based on a non-parametric approach that uses several visual features to calculate distances between image samples. While this approach is successful from the viewpoint of annotation accuracy, the computational costs, in terms of both complexity and memory use, tend to be high, since non-parametric methods require many training instances to be stored in memory to compute distances from a query. In this paper, we investigate several linear dimensionality reduction methods for efficient image annotation. Using the additional information provided by multiple labels, we can obtain a small representation preserving (and hopefully improving) the semantic distance of a visual feature. Linear methods are computationally reasonable and are suitable for practical large-scale systems, although only limited comparison of such methods is available in this research field. Extensive experiments and analyses on various datasets and visual features show how these simple methods can be applied effectively to image annotation.
Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
BMVC1
2010 Global Gaussian approach for scene categorization using information geometry
abstract
Local features provide powerful cues for generic image recognition. An image is represented by a “bag” of local features, which form a probabilistic distribution in the feature space. The problem is how to exploit the distributions efficiently. One of the most successful approaches is the bag-of-keypoints scheme, which can be interpreted as sparse sampling of high-level statistics, in the sense that it describes a complex structure of a local feature distribution using a relatively small number of parameters. In this paper, we propose the opposite approach, dense sampling of low-level statistics. A distribution is represented by a Gaussian in the entire feature space. We define some similarity measures of the distributions based on an information geometry framework and show how this conceptually simple approach can provide a satisfactory performance, comparable to the bag-of-keypoints for scene classification tasks. Furthermore, because our method and bag-of-keypoints illustrate different statistical points, we can further improve classification performance by using both of them in kernels.
Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
CVPR1
2010 Improving Local Descriptors by Embedding Global and Local Spatial Information
Tatsuya Harada, Hideki Nakayama, Yasuo Kuniyoshi
ECCV (4)2
2010 High-speed 3D object recognition using additive features in a linear subspace
abstract
In this paper we propose a method of high-speed 3D object recognition using linear subspace method and our 3D features. This method can be applied to partial models with any size in any posture. Although it is becoming easy to obtain textured 3D models by a 3D scanner, there are few methods for 3D object recognition which take into account both shape and textures of objects. Moreover, it is difficult to achieve high-speed processing of large 3D data. Our 3D features consider the co-occurrence of shape and colors of an object's surface. The additive property of these features makes it possible to calculate the similarity between a query part and the subspace of each object in a database without division, and therefore the time for recognition is quite short. In the experiments, we compare our method with conventional methods using Spin-Images and Textured Spin-Images. We show that our method is appropriate for 3D object recognition.
Asako Kanezaki, Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
ICRA2
2009 Image annotation and retrieval based on efficient learning of contextual latent space
abstract
Image annotation and retrieval are extremely difficult because of the generic nature of the target images. Generic images contain various miscellaneous objects and scenes. Therefore, desirable annotation results are subjective and underspecified. To overcome this problem, it is important to assume "weak labeling" framework, where images are weakly related to multiple words without region information. In this paper, we propose a high speed and high accuracy image annotation and retrieval method based on efficient learning of the contextual latent space. A distance between samples can be defined in the intrinsic feature space for annotation using latent space learning between images and labels. The proposed method is shown to be faster and more accurate than previously published methods.
Tatsuya Harada, Hideki Nakayama, Yasuo Kuniyoshi
ICME2
2007 Journalist robot: robot system making news articles from real world
abstract
We describe the development of a journalist robot system, which generates articles by searching for news in the real world. Our system repeats the steps: (1) autonomous exploration (2) recording of news, and (3) generation of articles. We characterize events with two values: "anomaly" and "relevance" to the user. During the exploration step, images are evaluated using these values. If an interesting event is detected, the robot approaches it to collect additional information. The system then labels the images, and generates a description from the labels. Experiments show the ability of our system to find news-like phenomena and describe images with words.
Rie Matsumoto, Hideki Nakayama, Tatsuya Harada, Yasuo Kuniyoshi
IROS2
2003 Towards Implicit Invocation of Web Services Functions
Takehiro Tokuda, Tetsuya Suzuki, Hideki Nakayama
EJC3