EDBT 2026 Demo / reviewers in the wild / expert
Nakamasa Inoue
dblp:27/8618
· DBLP profile ↗
74ranked-venue papers
17as first author
50since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 15 first-author · 41 since 2021Artificial intelligence and machine learning · 41 · 7 first-author · 29 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image CaptioningabstractLarge vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decoder (DISCODE), a novel finetuning-free method that generates robust evaluation scores better aligned with human judgments across diverse domains. The core idea behind DISCODE lies in its test-time adaptive evaluation approach, which introduces the Adaptive Test-Time (ATT) loss, leveraging a Gaussian prior distribution to improve robustness in evaluation score estimation. This loss is efficiently minimized at test time using an analytical solution that we derive. Furthermore, we introduce the Multi-domain Caption Evaluation (MCEval) benchmark, a new image captioning evaluation benchmark covering six distinct domains, designed to assess the robustness of evaluation metrics. In our experiments, we demonstrate that DISCODE achieves state-of-the-art performance as a reference-free evaluation metric across MCEval and four representative existing benchmarks. Nakamasa Inoue, Kanoko Goto, Masanari Oi, Martyna Gruszka, Mahiro Ukai, Takumi Hirose, Yusuke Sekikawa |
AAAI | 1 |
| 2026 | Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
Masanari Oi, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue |
LREC | 4 |
| 2026 | DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in InteractionsabstractModeling daily hand interactions often struggles with severe occlusions, such as when two hands overlap, which highlights the need for robust feature learning in 3D hand pose estimation (HPE). To handle such occluded hand images, it is vital to effectively learn the relationship between local image features (e.g., for occluded joints) and global context (e.g., cues from inter-joints, inter-hands, or the scene). However, most current 3D HPE methods still rely on ResNet for feature extraction, and such CNN’s inductive bias may not be optimal for 3D HPE due to its limited capability to model the global context. To address this limitation, we propose an effective and efficient framework for visual feature extraction in 3D HPE using recent state space modeling (i.e., Mamba), dubbed Deformable Mamba (DF-Mamba). DF-Mamba is designed to capture global context cues beyond standard convolution through Mamba’s selective state modeling and the proposed deformable state scanning. Specifically, for local features after convolution, our deformable scanning aggregates these features within an image while selectively preserving useful cues that represent the global context. This approach significantly improves the accuracy of structured 3D HPE, with comparable inference speed to ResNet-50. Our experiments involve extensive evaluations on five divergent datasets including single-hand and two-hand scenarios, hand-only and hand-object interactions, as well as RGB and depth-based estimation. DF-Mamba outperforms the latest image backbones, including VMamba and Spatial-Mamba, on all datasets and achieves state-of-the-art performance. Takehiko Ohkawa, Guwenxiao Zhou, Kanoko Goto, Takumi Hirose, Yusuke Sekikawa, Nakamasa Inoue |
WACV | 7 |
| 2025 | Rectified Lagrangian for Out-of-Distribution Detection in Modern Hopfield NetworksabstractModern Hopfield networks (MHNs) have recently gained significant attention in the field of artificial intelligence because they can store and retrieve a large set of patterns with an exponentially large memory capacity. A MHN is generally a dynamical system defined with Lagrangians of memory and feature neurons,where memories associated with in-distribution (ID) samples are represented by attractors in the feature space. One major problem in existing MHNs lies in managing out-of-distribution (OOD) samples because it was originally assumed that all samples are ID samples. To address this, we propose the rectified Lagrangian (RegLag), a new Lagrangian for memory neurons that explicitly incorporates an attractor for OOD samples in the dynamical system of MHNs. RecLag creates a trivial point attractor for any interaction matrix, enabling OOD detection by identifying samples that fall into this attractor as OOD. The interaction matrix is optimized so that the probability densities can be estimated to identify ID/OOD. We demonstrate the effectiveness of RecLag-based MHNs compared to energy-based OOD detection methods, including those using state-of-the-art Hopfield energies, across nine image datasets. Ryo Moriai, Nakamasa Inoue, Masayuki Tanaka 0001, Rei Kawakami, Satoshi Ikehata, Ikuro Sato |
AAAI | 2 |
| 2025 | Multi-Point Positional Insertion Tuning for Small Object DetectionabstractSmall object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensive. To address this issue, this paper introduces multi-point positional insertion (MPI) tuning, a parameter-efficient finetuning (PEFT) method for small object detection. Specifically, MPI incorporates multiple positional embeddings into a frozen pretrained model, enabling the efficient detection of small objects by providing precise positional information to latent features. Through experiments, we demonstrated the effectiveness of the proposed method on the SODA-D dataset. MPI performed comparably to conventional PEFT methods, including CoOp and VPT, while significantly reducing the number of parameters that need to be tuned. Kanoko Goto, Takumi Karasawa, Takumi Hirose, Rei Kawakami, Nakamasa Inoue |
ICASSP | 5 |
| 2025 | Binary Stochastic Flip Optimization for Training Binary Neural NetworksabstractFor deploying deep neural networks on edge devices with limited resources, binary neural networks (BNNs) have attracted significant attention, due to their computational and memory efficiency. However, once a neural network is binarized, finetuning it on edge devices becomes challenging because most conventional training algorithms for BNNs are designed for use on centralized servers and require storing real-valued parameters during training. To address this limitation, this paper introduces binary stochastic flip optimization (BinSFO), a novel training algorithm for BNNs. BinSFO employs a parameter update rule based on Boolean operations, eliminating the need to store real-valued parameters and thereby reducing memory requirements and computational overhead. In experiments, we demonstrated the effectiveness and memory efficiency of BinSFO in fine-tuning scenarios on six image classification datasets. BinSFO performed comparably to conventional training algorithms with a 70.7% smaller memory requirement. Code is released at https://github.com/TatsukichiShibuya/ICASSP2025_BinSFO Tatsukichi Shibuya, Nakamasa Inoue, Rei Kawakami, Ikuro Sato |
ICASSP | 2 |
| 2025 | Referring Expression Comprehension for Small ObjectsabstractReferring expression comprehension (REC) aims to localize the target object described by a natural language expression. Recent advances in vision-language learning have led to significant performance improvements in REC tasks. However, localizing extremely small objects remains a considerable challenge despite its importance in real-world applications such as autonomous driving. To address this issue, we introduce a novel dataset and method for REC targeting small objects. First, we present the small object REC (SOREC) dataset, which consists of 100,000 pairs of referring expressions and corresponding bounding boxes for small objects in driving scenarios. Second, we propose the progressive-iterative zooming adapter (PIZA), an adapter module for parameter-efficient fine-tuning that enables models to progressively zoom in and localize small objects. In a series of experiments, we apply PIZA to GroundingDINO and demonstrate a significant improvement in accuracy on the SOREC dataset. Our dataset, codes and pre-trained models are publicly available on the project page. Kanoko Goto, Takumi Hirose, Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue |
ICCV | 5 |
| 2025 | CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue |
ICCV | 7 |
| 2025 | AgroBench: Vision-Language Model Benchmark in Agriculture
Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi, Yoshitaka Ushiku |
ICCV | 2 |
| 2025 | AnimalClue: Recognizing Animals by their Traces
Risa Shinoda, Nakamasa Inoue, Iro Laina, Christian Rupprecht 0001, Hirokatsu Kataoka |
ICCV | 2 |
| 2025 | GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsabstractThe advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings. To overcome these limitations, we propose GeoProg3D, a visual programming framework that enables natural language-driven interactions with city-scale high-fidelity 3D scenes. GeoProg3D consists of two key components: (i) a Geography-aware City-scale 3D Language Field (GCLF) that leverages a memory-efficient hierarchical 3D model to handle large-scale data, integrated with geographic information for efficiently filtering vast urban spaces using directional cues, distance measurements, elevation data, and landmark references; and (ii) Geographical Vision APIs (GV-APIs), specialized geographic vision tools such as area segmentation and object detection. Our framework employs large language models (LLMs) as reasoning engines to dynamically combine GV-APIs and operate GCLF, effectively supporting diverse geographic vision tasks. To assess performance in city-scale reasoning, we introduce GeoEval3D, a comprehensive benchmark dataset containing 952 query-answer pairs across five challenging tasks: grounding, spatial reasoning, comparison, counting, and measurement. Experiments demonstrate that GeoProg3D significantly outperforms existing 3D language fields and vision-language models across multiple tasks. To our knowledge, GeoProg3D is the first framework enabling compositional geographic reasoning in high-fidelity city-scale 3D environments via natural language. The code is available at https://snskysk.github.io/GeoProg3D/. Shunsuke Yasuki, Taiki Miyanishi, Nakamasa Inoue, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Masato Taki, Yutaka Matsuo |
ICCV | 3 |
| 2025 | HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech SynthesisabstractRecently, Text-to-speech (TTS) models based on large language models (LLMs)
that translate natural language text into sequences of discrete audio tokens have
gained great research attention, with advances in neural audio codec (NAC) mod-
els using residual vector quantization (RVQ). However, long-form speech synthe-
sis remains a significant challenge due to the high frame rate, which increases the
length of audio tokens and makes it difficult for autoregressive language models
to generate audio tokens for even a minute of speech. To address this challenge,
this paper introduces two novel post-training approaches: 1) Multi-Resolution Re-
quantization (MReQ) and 2) HALL-E. MReQ is a framework to reduce the frame
rate of pre-trained NAC models. Specifically, it incorporates multi-resolution
residual vector quantization (MRVQ) module that hierarchically reorganizes dis-
crete audio tokens through teacher-student distillation. HALL-E is an LLM-based
TTS model designed to predict hierarchical tokens of MReQ. Specifically, it incor-
porates the technique of using MRVQ sub-modules and continues training from a
pre-trained LLM-based TTS model. Furthermore, to promote TTS research, we
create MinutesSpeech, a new benchmark dataset consisting of 40k hours of filtered
speech data for training and evaluating speech synthesis ranging from 3s up to
180s. In experiments, we demonstrated the effectiveness of our approaches by ap-
plying our post-training framework to VALL-E. We achieved the frame rate down
to as low as 8 Hz, enabling the stable minitue-long speech synthesis in a single
inference step. Audio samples, dataset, codes and pre-trained models are available
at https://yutonishimura-v2.github.io/HALL-E_DEMO. Yuto Nishimura, Takumi Hirose, Masanari Ohi, Hideki Nakayama, Nakamasa Inoue |
ICLR | 5 |
| 2025 | Collision Avoidance with Differentiable Occupancy Functions in Object RearrangementabstractWe address the challenge of object relocation by robots in environments where their behavior is expected to resemble that of humans. Existing methods typically learn to regress the position and orientation of objects specified by natural language commands using training data. However, these approaches do not account for physical constraints during training, often resulting in collisions between relocated objects. In this work, we introduce a collision avoidance loss based on functions that incorporate object size into the training process. Specifically, we propose a type of occupancy function in which particles are represented by a 3D Gaussian probability density function. By incorporating these functions into an additional training phase of existing models, we demonstrate a reduction in the number of collisions during rearrangement tasks. Notably, despite the decrease in collisions, the semantic structure of the relocation results is preserved. Roma Satoh, Nakamasa Inoue, Rei Kawakami |
IROS | 2 |
| 2025 | STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language ModelsabstractObject state recognition aims to identify the specific condition of objects, such as their positional states (e.g., open or closed) and functional states (e.g., on or off). While recent Vision-Language Models (VLMs) are capable of performing a variety of multimodal tasks, it remains unclear how precisely they can identify object states. To alleviate this issue, we introduce the STAte and Transition UnderStanding Benchmark (STATUS Bench), the first benchmark for rigorously evaluating the ability of VLMs to understand subtle variations in object states in diverse situations. Specifically, STATUS Bench introduces a novel evaluation scheme that requires VLMs to perform three tasks simultaneously: object state identification (OSI), image retrieval (IR), and state change identification (SCI). These tasks are defined over our fully hand-crafted dataset involving image pairs, their corresponding object state descriptions and state change descriptions. Furthermore, we introduce a large-scale training dataset, namely STATUS Train, which consists of 13 million semi-automatically created descriptions. This dataset serves as the largest resource to facilitate further research in this area. In our experiments, we demonstrate that STATUS Bench enables rigorous consistency evaluation and reveal that current state-of-the-art VLMs still significantly struggle to capture subtle object state distinctions. Surprisingly, under the proposed rigorous evaluation scheme, most open-weight VLMs exhibited chance-level zero-shot performance. After fine-tuning on STATUS Train, Qwen2.5-VL achieved performance comparable to Gemini 2.0 Flash. These findings underscore the necessity of STATUS Bench and Train for advancing object state recognition in VLM research. Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue |
ACM Multimedia | 3 |
| 2025 | Contracted Gram Tensor Distillation for Object DetectionabstractKnowledge distillation (KD) is a technique for compressing large models into lightweight ones while preserving performance and has also proven effective in object detection. However, most existing detection KD methods naively adapt techniques designed for classification tasks, overlooking the uniqueness of the task output space in detection, including batch, class, and spatial dimensions. To address this limitation, we propose Contracted Gram Tensor Distillation (CGTD), a novel knowledge distillation framework for object detection. CGTD effectively distills structural knowledge in high-order tensors by minimizing loss between contracted Gram tensors, each of which is calculated along each axis of the target tensors and efficiently captures structural information. This enables the student model to learn not just the output values but also the underlying structural patterns, maximizing the effectiveness of knowledge distillation in object detection. We conduct experiments on the COCO and SODA-D datasets and show that our framework achieves state-of-the-art results in object detection distillation. Takumi Karasawa, Nakamasa Inoue, Rei Kawakami |
MMAsia | 2 |
| 2025 | Masked Gated Linear UnitabstractGated Linear Units (GLUs) have become essential components in the feed-forward networks of state-of-the-art Large Language Models (LLMs).
However, they require twice as many memory reads compared to feed-forward layers without gating, due to the use of separate weight matrices for the gate and value streams.
To address this bottleneck, we introduce Masked Gated Linear Units (MGLUs), a novel family of GLUs with an efficient kernel implementation.
The core contribution of MGLUs include:
(1) the Mixture of Element-wise Gating (MoEG) architecture that learns multiple binary masks, each determining gate or value assignments at the element level on a single shared weight matrix resulting in reduced memory transfer, and (2) FlashMGLU, a hardware-friendly kernel that yields up to a 19.7$\times$ inference-time speed-up over a na\"ive PyTorch MGLU and is 47\% more memory-efficient and 34\% faster than standard GLUs despite added architectural complexity on an RTX5090 GPU.
In LLM experiments, the Swish-activated variant SwiMGLU preserves its memory advantages while matching—or even surpassing—the downstream accuracy of the SwiGLU baseline. Yukito Tajima, Nakamasa Inoue, Yusuke Sekikawa, Ikuro Sato, Rio Yokota |
NeurIPS | 2 |
| 2025 | Diffusion-Based Generative Regularization for Supervised Discriminative LearningabstractEnsuring the quality and quantity of labeled training data has long been a challenge in training deep neural networks for discriminative tasks. One solution to this problem is to use a generative model to augment training data and learn a discriminative model with it. For image classification, with the recent development of diffusion models, it has become possible to generate a variety synthetic images, and there are high expectations for their use as training data. However, to obtain high-quality labeled synthetic images, the hyperparameters and prompts often need to be manually tuned, and the accuracy of the trained image classification model is highly dependent on them. To address this issue, this paper proposes diffusion-based generative regularization, a supervised discriminative learning framework that utilizes a diffusion-based image generation model as a regularizer to robustly learn discriminative representations without the need to synthesize images. Our experiments using vision transformers and stable diffusion models on ImageNet-1k demonstrate that the proposed framework improves classification accuracy on both in-distribution and distribution-shifted data. Takuya Asakura, Nakamasa Inoue, Koichi Shinoda |
WACV | 2 |
| 2025 | ContextualCoder: Adaptive In-Context Prompting for Programmatic Visual Question AnsweringabstractVisual Question Answering (VQA) presents a challenging task at the intersection of computer vision and natural language processing, aiming to bridge the semantic gap between visual perception and linguistic comprehension. Traditional VQA approaches do not distinguish between data processing and reasoning, limiting their interpretability and generalizability in complex and diverse scenarios. Conversely, Programmatic Visual Question Answering (PVQA) models leverage large language models (LLMs) to generate executable codes, providing answers with detailed and interpretable reasoning processes. However, existing PVQA models typically rely on simplistic input-output prompting, which struggles to elicit domain-specific knowledge from LLMs and often produces unclear or extraneous outputs. Furthermore, PVQA models typically rely on a basic in-context example (ICE) selection methodology that is heavily influenced by individual word similarity rather than the overall sentence context. This leads to suboptimal ICE selection and a reliance on dataset-specific ICE candidates. In this paper, we propose ContextualCoder, a novel prompting framework tailored for PVQA models. ContextualCoder leverages frozen LLMs for code generation and pre-trained visual models for code execution, eliminating the need for extensive training and enhancing model flexibility. By incorporating an innovative prompting methodology and a novel ICE selection strategy, ContextualCoder facilitates the use of diverse in-context information for code generation, thereby improving the performance of PVQA models. Our approach surpasses state-of-the-art models, as evidenced by comprehensive experiments across diverse VQA datasets, including multilingual scenarios. Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex Chichung Kot, Koichi Shinoda |
IEEE Trans. Multim. | 2 |
| 2024 | Efficient Target Propagation by Deriving Analytical SolutionabstractExploring biologically plausible algorithms as alternatives to error backpropagation (BP) is a challenging research topic in artificial intelligence. It also provides insights into the brain's learning methods. Recently, when combined with well-designed feedback loss functions such as Local Difference Reconstruction Loss (LDRL) and through hierarchical training of feedback pathway synaptic weights, Target Propagation (TP) has achieved performance comparable to BP in image classification tasks. However, with an increase in the number of network layers, the tuning and training cost of feedback weights escalates. Drawing inspiration from the work of Ernoult et al., we propose a training method that seeks the optimal solution for feedback weights. This method enhances the efficiency of feedback training by analytically minimizing feedback loss, allowing the feedback layer to skip certain local training iterations. More specifically, we introduce the Jacobian matching loss (JML) for feedback training. We also proactively implement layers designed to derive analytical solutions that minimize JML. Through experiments, we have validated the effectiveness of this approach. Using the CIFAR-10 dataset, our method showcases accuracy levels comparable to state-of-the-art TP methods. Furthermore, we have explored its effectiveness in more intricate network architectures. Yanhao Bao, Tatsukichi Shibuya, Ikuro Sato, Rei Kawakami, Nakamasa Inoue |
AAAI | 5 |
| 2024 | A Simple Finetuning Strategy Based on Bias-Variance Ratios of Layer-Wise Gradients
Mao Tomita, Ikuro Sato, Rei Kawakami, Nakamasa Inoue, Satoshi Ikehata, Masayuki Tanaka 0001 |
ACCV (8) | 4 |
| 2024 | Scaling Backwards: Minimal Synthetic Pre-Training?
Ryu Tadokoro, Ryosuke Yamada, Yuki Markus Asano, Iro Laina, Christian Rupprecht 0001, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka |
ECCV (15) | 7 |
| 2024 | Rethinking Image Super-Resolution from Training Data Perspectives
Go Ohtani, Ryu Tadokoro, Ryosuke Yamada, Yuki Markus Asano, Iro Laina, Christian Rupprecht 0001, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka, Yoshimitsu Aoki |
ECCV (17) | 7 |
| 2024 | Formula-Supervised Visual-Geometric Pre-training
Ryosuke Yamada, Kensho Hara, Hirokatsu Kataoka, Koshi Makihara, Nakamasa Inoue, Rio Yokota, Yutaka Satoh |
ECCV (22) | 5 |
| 2024 | Cubic Knowledge Distillation for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) can play an important role in human-computer interaction. In this paper, we propose a logit knowledge distillation method for SER, called Cubic KD, that distill the knowledge of fine-tuned self-supervised models to allow better performance of small models. By creating cubic structures from teacher and student network output features and using a loss function to distill the cube structure through self-correlation between elements, Cubic KD efficiently captures knowledge within instances and among instances. We apply this distillation method to four student models and conduct experiments using the Emo-DB and IEMOCAP datasets. The results show that Cubic KD outperforms existing predictive logit knowledge distillation methods and is comparable to intermediate feature knowledge distillation methods. Our implementation code is available at https://github.com/Fly1toMoon/Cubic-Knowledge-Distillation Zhibo Lou, Shinta Otake, Zhengxiao Li, Rei Kawakami, Nakamasa Inoue |
ICASSP | 5 |
| 2024 | PolarDB: Formula-Driven Dataset for Pre-Training Trajectory EncodersabstractFormula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bias and racial bias because it does not rely on real data as discussed in previous studies using fractals and polygons for pre-training image encoders. While FDSL has been proposed for pre-training image encoders, it has not been considered for temporal trajectory data. In this paper, we introduce PolarDB, the first formula-driven dataset for pre-training trajectory encoders with an application to fine-grained cutting-method recognition using hand trajectories. More specifically, we generate 270k trajectories for 432 categories on the basis of polar equations and use them to pre-train a Transformer-based trajectory encoder in an FDSL manner. In the experiments, we show that pre-training on PolarDB improves the accuracy of fine-grained cutting-method recognition on cooking videos of EPIC-KITCHEN and Ego4D datasets, where the pre-trained trajectory encoder is used as a plug-in module for a video recognition network. Sota Miyamoto, Takuma Yagi, Yuto Makimoto, Mahiro Ukai, Yoshitaka Ushiku, Atsushi Hashimoto 0001, Nakamasa Inoue |
ICASSP | 7 |
| 2024 | Pseudo-Outlier Synthesis Using Q-Gaussian Distributions for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection, which aims to determine whether an input is outside the training data distribution or not, is an indispensable task in many computer vision applications. In many of the previous studies on OOD detection for image classification, the class-conditional distribution of visual features is assumed to be a Gaussian. However, this may not be a reasonable assumption because unseen outliers do not always follow a Gaussian distribution. In this study, we investigated the potential effects of non-Gaussian distributions by using an OOD detection method based on Tsallis statistics, in which the family of q-Gaussian distributions involving short- and long-tail distributions are used for synthesizing pseudo outlier features for improving the effectiveness of training. In experiments on six image classification datasets, we show that the proposed method achieves good results in the comparison method. In addition, we find that samples with a smaller hem than the Gaussian distribution by all datasets by ablation studies of the tail of the distribution improve the performance of OOD detection. Ryu Tadokoro, Eisuke Yamagata, Yusuke Kondo, Kensho Hara, Hirokatsu Kataoka, Nakamasa Inoue |
ICASSP | 7 |
| 2024 | Spatiality-Aware Prompt Tuning for Few-Shot Small Object DetectionabstractSmall Object Detection (SOD) is challenging due to the scarcity of image features arising from small image regions. The niche nature of small objects additionally poses difficulty in data collection compared to normal-sized objects. Therefore, efficient learning from limited data is benefical for SOD. To tackle few-shot SOD, we propose Spatiality-Aware Prompt Tuning (SAPT), a novel prompt tuning method for vision-language models (VLMs) to deal with the image feature scarcity and the limited data for small objects. SAPT appends the verbalizer prompt, expressing the spatiality of small objects through a template-based sentence, to the text prompt of the pre-trained VLMs. During fine-tuning, the integrated text prompt is learned solely by the decoder of the vision-language detector, while the image and text backbones of the model remain frozen to facilitate efficient learning. In our experiments, we demonstrate the effectiveness of the proposed method on SODA-D and COCO datasets in few-shot and full-shot learning scenarios, and show that our method improves state-of-the-art in both scenarios. Takumi Karasawa, Nakamasa Inoue, Rei Kawakami |
ICIP | 2 |
| 2024 | Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question AnsweringabstractVisual question answering (VQA) is the task of providing accurate answers to natural language questions based on visual input. Programmatic VQA (PVQA) models have been gaining attention recently. These use large language models (LLMs) to formulate executable programs that address questions requiring complex visual reasoning. However, there are challenges in enabling LLMs to comprehend the usage of image processing modules and generate relevant code. To overcome these challenges, this paper introduces PyramidCoder, a novel prompting framework for PVQA models. PyramidCoder consists of three hierarchical levels, each serving a distinct purpose: query rephrasing, code generation, and answer aggregation. Notably, PyramidCoder utilizes a single frozen LLM and pre-defined prompts at each level, eliminating the need for additional training and ensuring flexibility across various LLM architectures. Compared to the state-of-the-art PVQA model, our approach improves accuracy by at least 0.5% on the GQA dataset, 1.4% on the VQAv2 dataset, and 2.9% on the NLVR2 dataset. Ruoyue Shen, Nakamasa Inoue, Koichi Shinoda |
ICIP | 2 |
| 2024 | On the Relationship Between Double Descent of CNNs and Shape/Texture Bias Under Learning Process
Shun Iwase, Shuya Takahashi, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka, Eisaku Maeda |
ICPR (25) | 3 |
| 2024 | Locally Aligned Rectified Flow Model for Speech Enhancement Towards Single-Step Diffusion
Zhengxiao Li, Nakamasa Inoue |
INTERSPEECH | 2 |
| 2024 | AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question AnsweringabstractVisual question answering aims to provide responses to questions given visual input. Recently, visual programmatic models (VPMs), which generate programs to answer questions through large language models (LLMs), have attracted attention. However, they often require long input prompts to provide the LLM with sufficient API usage details to generate relevant code. To address this limitation, we propose AdaCoder, an adaptive prompt compression framework for VPMs. AdaCoder operates in two phases: a compression phase and an inference phase. In the compression phase, given a preprompt that describes all API definitions with example code snippets, a set of compressed preprompts is generated, each depending on a specific question type. In the inference phase, AdaCoder predicts the question type and chooses the appropriate corresponding compressed preprompt to generate code to answer the question. In experiments, we apply AdaCoder to ViperGPT and demonstrate that it reduces token length by 71.1%, while maintaining or even improving the performance of visual question answering. Mahiro Ukai, Shuhei Kurita, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Nakamasa Inoue |
ACM Multimedia | 5 |
| 2024 | ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing TasksabstractSelf-supervised learning has emerged as a key approach for learning generic representations from speech data. Despite promising results in downstream tasks such as speech recognition, speaker verification, and emotion recognition, a significant number of parameters is required, which makes fine-tuning for each task memory-inefficient. To address this limitation, we introduce ELP-adapter tuning, a novel method for parameter-efficient fine-tuning using three types of adapter, namely encoder adapters (E-adapters), layer adapters (L-adapters), and a prompt adapter (P-adapter). The E-adapters are integrated into transformer-based encoder layers and help to learn finegrained speech representations that are effective for speech recognition. The L-adapters create paths from each encoder layer to the downstream head and help to extract non-linguistic features from lower encoder layers that are effective for speaker verification and emotion recognition. The P-adapter appends pseudo features to CNN features to further improve effectiveness and efficiency. With these adapters, models can be quickly adapted to various speech processing tasks. Our evaluation across four downstream tasks using five backbone models demonstrated the effectiveness of the proposed method. With the WavLM backbone, its performance was comparable to or better than that of full fine-tuning on all tasks while requiring 90% fewer learnable parameters. Nakamasa Inoue, Shinta Otake, Takumi Hirose, Masanari Ohi, Rei Kawakami |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Fixed-Weight Difference Target PropagationabstractTarget Propagation (TP) is a biologically more plausible algorithm than the error backpropagation (BP) to train deep networks, and improving practicality of TP is an open issue. TP methods require the feedforward and feedback networks to form layer-wise autoencoders for propagating the target values generated at the output layer. However, this causes certain drawbacks; e.g., careful hyperparameter tuning is required to synchronize the feedforward and feedback training, and frequent updates of the feedback path are usually required than that of the feedforward path. Learning of the feedforward and feedback networks is sufficient to make TP methods capable of training, but is having these layer-wise autoencoders a necessary condition for TP to work? We answer this question by presenting Fixed-Weight Difference Target Propagation (FW-DTP) that keeps the feedback weights constant during training. We confirmed that this simple method, which naturally resolves the abovementioned problems of TP, can still deliver informative target values to hidden layers for a given task; indeed, FW-DTP consistently achieves higher test performance than a baseline, the Difference Target Propagation (DTP), on four classification datasets. We also present a novel propagation architecture that explains the exact form of the feedback function of DTP to analyze FW-DTP. Our code is available at https://github.com/TatsukichiShibuya/Fixed-Weight-Difference-Target-Propagation. Tatsukichi Shibuya, Nakamasa Inoue, Rei Kawakami, Ikuro Sato |
AAAI | 2 |
| 2023 | Learning with Partial Forgetting in Modern Hopfield NetworksabstractIt has been known by neuroscience studies that partial and transient forgetting of memory often plays an important role in the brain to improve performance for certain intellectual activities. In machine learning, associative memory models such as classical and modern Hopfield networks have been proposed to express memories as attractors in the feature space of a closed recurrent network. In this work, we propose learning with partial forgetting (LwPF), where a partial forgetting functionality is designed by element-wise non-bijective projections, for memory neurons in modern Hopfield networks to improve model performance. We incorporate LwPF into the attention mechanism also, whose process has been shown to be identical to the update rule of a certain modern Hopfield network, by modifying the corresponding Lagrangian. We evaluated the effectiveness of LwPF on three diverse tasks such as bit-pattern classification, immune repertoire classification for computational biology, and image classification for computer vision, and confirmed that LwPF consistently improves the performance of existing neural networks including DeepRC and vision transformers. Toshihiro Ota, Ikuro Sato, Rei Kawakami, Masayuki Tanaka 0001, Nakamasa Inoue |
AISTATS | 5 |
| 2023 | Visual Atoms: Pre-Training Vision Transformers with Sinusoidal WavesabstractFormula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision transformers. However, the lack of a systematic investigation as to why these contour-oriented synthetic datasets can achieve the same accuracy as real datasets leaves much room for skepticism. In the present work, we develop a novel methodology based on circular harmonics for systematically investigating the design space of contour-oriented synthetic datasets. This allows us to efficiently search the optimal range of FDSL parameters and maximize the variety of synthetic images in the dataset, which we found to be a critical factor. When the resulting new dataset VisualAtom-21k is used for pre-training ViT-Base, the top-1 accuracy reached 83.7% when fine-tuning on ImageNet-1k. This is only 0.5% difference from the top-1 accuracy (84.2%) achieved by the JFT-300M pre-training, even though the scale of images is 1/14. Unlike JFT-300M which is a static dataset, the quality of synthetic datasets will continue to improve, and the current work is a testament to this possibility. FDSL is also free of the common issues associated with real images, e.g. privacy/copyright issues, labeling costs/errors, and ethical biases. Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka, Rio Yokota |
CVPR | 3 |
| 2023 | Step restriction for improving adversarial attacksabstractWe propose an algorithm to automatically restrict the step size in the iterative optimization process with an application to adversarial attacks on speaker verification models. The proposed algorithm dynamically determines a subspace with a restriction radius r to which the Taylor approximation is applied at each iteration and then solves a linear problem within the subspace by using the projected gradient method. In experiments, we demonstrate adversarial attacks on three speaker verification models: i-vectors, SE-ResNet-34, and ECAPATDNN. We show that the degree of adversarial perturbations generated by the proposed algorithm is smaller than that generated by the conventional attack method. Keita Goto, Shinta Otake, Rei Kawakami, Nakamasa Inoue |
ICASSP | 4 |
| 2023 | Parameter Efficient Transfer Learning for Various Speech Processing TasksabstractFine-tuning of self-supervised models is a powerful transfer learning method in a variety of fields, including speech processing, since it can utilize generic feature representations obtained from large amounts of unlabeled data. Fine-tuning, however, requires a new parameter set for each downstream task, which is parameter inefficient. Adapter architecture is proposed to partially solve this issue by inserting lightweight learnable modules into a frozen pre-trained model. However, existing adapter architectures fail to adaptively leverage low-to high-level features stored in different layers, which is necessary for solving various kinds of speech processing tasks. Thus, we propose a new adapter architecture to acquire feature representations more flexibly for various speech tasks. In experiments, we applied this adapter to WavLM on four speech tasks. It performed on par or better than naïve fine-tuning, with only 11% of learnable parameters. It also outperformed an existing adapter architecture. Our implementation code is available at https://github.com/sinhat98/adapter-wavlm Shinta Otake, Rei Kawakami, Nakamasa Inoue |
ICASSP | 3 |
| 2023 | Pre-training Vision Transformers with Very Limited Synthesized ImagesabstractFormula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers on such synthetic datasets can yield competitive accuracy on a wide range of downstream tasks. These synthetic images are categorized according to the parameters in the mathematical formula that generate them. In the present work, we hypothesize that the process for generating different instances for the same category in FDSL, can be viewed as a form of data augmentation. We validate this hypothesis by replacing the instances with data augmentation, which means we only need a single image per category. Our experiments shows that this one-instance fractal database (OFDB) performs better than the original dataset where instances were explicitly generated. We further scale up OFDB to 21,000 categories and show that it matches, or even surpasses, the model pre-trained on ImageNet-21k in ImageNet-1k fine-tuning. The number of images in OFDB is 21k, whereas ImageNet-21k has 14M. This opens new possibilities for pre-training vision transformers with much smaller datasets. Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega, Rio Yokota, Nakamasa Inoue |
ICCV | 6 |
| 2023 | SegRCDB: Semantic Segmentation via Formula-Driven Supervised LearningabstractPre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels is quite difficult to construct. Moreover, what matters in semantic segmentation pre-training has not been fully investigated. In this paper, we propose the Segmentation Radial Contour DataBase (SegRCDB), which for the first time applies formula-driven supervised learning for semantic segmentation. SegRCDB enables pre-training for semantic segmentation without real images or any manual semantic labels. SegRCDB is based on insights about what is important in pre-training for semantic segmentation and allows efficient pre-training. Pre-training with SegRCDB achieved higher mIoU than the pre-training with COCO-Stuff for fine-tuning on ADE-20k and Cityscapes with the same number of training images. SegRCDB has a high potential to contribute to semantic segmentation pre-training and investigation by enabling the creation of large datasets without manual annotation. The SegRCDB dataset will be released under a license that allows research and commercial use. Code is available at: https://github.com/dahlian00/SegRCDB Risa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka |
ICCV | 4 |
| 2023 | Scale-space Tokenization for Improving the Robustness of Vision TransformersabstractThe performance of the Vision Transformer (ViT) model and its variants in most vision tasks has surpassed traditional Convolutional Neural Networks (CNNs) in terms of in-distribution accuracy. However, ViTs still have significant room for improvement in their robustness to input perturbations. Furthermore, robustness is a critical aspect to consider when deploying ViTs in real-world scenarios. Despite this, some variants of ViT improve the in-distribution accuracy and computation performance at the cost of sacrificing the model's robustness and generalization. In this study, inspired by the prior findings on the potential effectiveness of shape bias to robustness improvement and the importance of multi-scale analysis, we propose a simple yet effective method, scale-space tokenization, to improve the robustness of ViT while maintaining in-distribution accuracy. Based on this method, we build Scale-space-based Robust Vision Transformer (SRVT) model. Our method consists of scale-space patch embedding and scale-space positional encoding. The scale-space patch embedding makes a sequence of variable-scale images and increases the model's shape bias to enhance its robustness. The scale-space positional encoding implicitly boosts the model's invariance to input perturbations by incorporating scale-aware position information into 3D sinusoidal positional encoding. We conduct experiments on image recognition benchmarks (CIFAR10/100 and ImageNet-1k) from the perspectives of in-distribution accuracy, adversarial and out-of-distribution robustness. The experimental results demonstrate our method's effectiveness in improving robustness without compromising in-distribution accuracy. Especially, our approach achieves advanced adversarial robustness on ImageNet-1k benchmark compared with state-of-the-art robust ViT. Rei Kawakami, Nakamasa Inoue |
ACM Multimedia | 3 |
| 2023 | CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud DataabstractCity-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings that can be utilized for attractive applications such as user-interactive navigation of autonomous vehicles and drones. However, compared to the extensive text annotations available for images and indoor scenes, the scarcity of text annotations for outdoor scenes poses a significant challenge for achieving these applications. To tackle this problem, we introduce the CityRefer dataset for city-level visual grounding. The dataset consists of 35k natural language descriptions of 3D objects appearing in SensatUrban city scenes and 5k landmarks labels synchronizing with OpenStreetMap. To ensure the quality and accuracy of the dataset, all descriptions and labels in the CityRefer dataset are manually verified. We also have developed a baseline system that can learn encoded language descriptions, 3D object instances, and geographical information about the city's landmarks to perform visual grounding on the CityRefer dataset. To the best of our knowledge, the CityRefer dataset is the largest city-level visual grounding dataset for localizing specific 3D objects. Taiki Miyanishi, Fumiya Kitamori, Shuhei Kurita, Jungdae Lee, Motoaki Kawanabe, Nakamasa Inoue |
NeurIPS | 6 |
| 2023 | Text-Guided Object Detector for Multi-modal Video Question AnsweringabstractVideo Question Answering (Video QA) is a task to answer a text-format question based on the understanding of linguistic semantics, visual information, and also linguistic-visual alignment in the video. In Video QA, an object detector pre-trained with large-scale datasets, such as Faster R-CNN, has been widely used to extract visual representations from video frames. However, it is not always able to precisely detect the objects needed to answer the question be-cause of the domain gaps between the datasets for training the object detector and those for Video QA. In this paper, we propose a text-guided object detector (TGOD), which takes text question-answer pairs and video frames as inputs, detects the objects relevant to the given text, and thus provides intuitive visualization and interpretable results. Our experiments using the STAGE framework on the TVQA+ dataset show the effectiveness of our proposed detector. It achieves a 2.02 points improvement in accuracy of QA, 12.13 points improvement in object detection (mAP50), 1.1 points improvement in temporal location, and 2.52 points improvement in ASA over the STAGE original detector. Ruoyue Shen, Nakamasa Inoue, Koichi Shinoda |
WACV | 2 |
| 2022 | Can Vision Transformers Learn without Natural Images?abstractIs it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated labels, the recent use of natural images has resulted in problems related to privacy violation, inadequate fairness protection, and the need for labor-intensive annotations. In this paper, we experimentally verify that the results of formula-driven supervised learning (FDSL) framework are comparable with, and can even partially outperform, sophisticated self-supervised learning (SSL) methods like SimCLRv2 and MoCov2 without using any natural images in the pre-training phase. We also consider ways to reorganize FractalDB generation based on our tentative conclusion that there is room for configuration improvements in the iterated function system (IFS) parameter settings of such databases. Moreover, we show that while ViTs pre-trained without natural images produce visualizations that are somewhat different from ImageNet pre-trained ViTs, they can still interpret natural image datasets to a large extent. Finally, in experiments using the CIFAR-10 dataset, we show that our model achieved a performance rate of 97.8, which is comparable to the rate of 97.4 achieved with SimCLRv2 and 98.0 achieved with ImageNet. Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, Yutaka Satoh |
AAAI | 5 |
| 2022 | Replacing Labeled Real-image Datasets with Auto-generated ContoursabstractIn the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageNet-21k shows 81.8% top-1 accuracy when fine-tuned on ImageNet-1k and FDSL shows 82.7% top-1 accuracy when pre-trained under the same conditions (number of images, hyperparameters, and number of epochs). Images generated by formulas avoid the privacy/copyright issues, labeling cost and errors, and biases that real images suffer from, and thus have tremendous potential for pre-training general models. To understand the performance of the synthetic images, we tested two hypotheses, namely (i) object contours are what matter in FDSL datasets and (ii) increased number of parameters to create labels affects performance improvement in FDSL pre-training. To test the former hypothesis, we constructed a dataset that consisted of simple object contour combinations. We found that this dataset can match the performance of fractals. For the latter hypothesis, we found that increasing the difficulty of the pre-training task generally leads to better fine-tuning accuracy. Hirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima, Sora Takashima, Edgar Josafat Martinez-Noriega, Nakamasa Inoue, Rio Yokota |
CVPR | 8 |
| 2022 | Downstream Augmentation Generation For Contrastive LearningabstractContrastive learning has become one of the most promising approaches for learning image representations. However, it heavily relies on heuristic data augmentation techniques, such as Gaussian blurring and color jittering, for making image pairs to be contrastively compared. These augmentations are not always appropriate for downstream tasks that each have their own camera and illumination settings. In this paper, we aim at improving the augmentation process and propose an augmentation generator, a network that learns to augment images for contrastive learning. Under the assumption that each downstream task has an optimal implicit augmentation function, the augmentation generator enhances the contrastive learning by estimating it. We demonstrate the effectiveness of our learning framework on two combined datasets, EMNIST-Omniglot and ImageNet-DAISO. Tomohiro Hayase, Suguru Yasutomi, Nakamasa Inoue |
ICASSP | 3 |
| 2022 | PoF: Post-Training of Feature Extractor for Improving GeneralizationabstractIt has been intensively investigated that the local shape, especially flatness, of the loss landscape near a minimum plays an important role for generalization of deep models. We developed a training algorithm called PoF: Post-Training of Feature Extractor that updates the feature extractor part of an already-trained deep model to search a flatter minimum. The characteristics are two-fold: 1) Feature extractor is trained under parameter perturbations in the higher-layer parameter space, based on observations that suggest flattening higher-layer parameter space, and 2) the perturbation range is determined in a data-driven manner aiming to reduce a part of test loss caused by the positive loss curvature. We provide a theoretical analysis that shows the proposed algorithm implicitly reduces the target Hessian components as well as the loss. Experimental results show that PoF improved model performance against baseline methods on both CIFAR-10 and CIFAR-100 datasets for only 10-epoch post-training, and on SVHN dataset for 50-epoch post-training. Ikuro Sato, Ryota Yamada, Masayuki Tanaka 0001, Nakamasa Inoue, Rei Kawakami |
ICML | 4 |
| 2022 | Spatiotemporal Initialization for 3D CNNs with Generated Motion PatternsabstractThe paper proposes a framework of Formula-Driven Supervised Learning (FDSL) for spatiotemporal initialization. Our FDSL approach enables to automatically and simultaneously generate motion patterns and their video labels with a simple formula which is based on Perlin noise. We designed a dataset of generated motion patterns adequate for the 3D CNNs to learn a better basis set of natural videos. The constructed Video Perlin Noise (VPN) dataset can be applied to initialize a model before pre-training with large-scale video datasets such as Kinetics-400/700, to enhance target task performance. Our spatiotemporal initialization with VPN dataset (VPN initialization) outperforms the previous initialization method with the inflated 3D ConvNet (I3D) using 2D ImageNet dataset. Our proposed method increased the top-1 video-level accuracy of Kinetics-400 pre-trained model on {Kinetics-400, UCF-101, HMDB-51, ActivityNet} datasets. Especially, the proposed method increased the performance rate of Kinetics-400 pre-trained model by 10.3 pt on ActivityNet. We also report that the relative performance improvements from the baseline are greater in 3D CNNs rather than other models. Our VPN initialization mainly helps to enhance the performance in spatiotemporal 3D kernels. The datasets, codes and pre-trained models used in this study will be publicly available1. Hirokatsu Kataoka, Kensho Hara, Ryusuke Hayashi, Eisuke Yamagata, Nakamasa Inoue |
WACV | 5 |
| 2022 | Pre-Training Without Natural ImagesabstractAbstract Is it possible to use convolutional neural networks pre-trained without any natural images to assist natural image understanding? The paper proposes a novel concept, Formula-driven Supervised Learning (FDSL). We automatically generate image patterns and their category labels by assigning fractals, which are based on a natural law. Theoretically, the use of automatically generated images instead of natural images in the pre-training phase allows us to generate an infinitely large dataset of labeled images. The proposed framework is similar yet different from Self-Supervised Learning because the FDSL framework enables the creation of image patterns based on any mathematical formulas in addition to self-generated labels. Further, unlike pre-training with a synthetic image dataset, a dataset under the framework of FDSL is not required to define object categories, surface texture, lighting conditions, and camera viewpoint. In the experimental section, we find a better dataset configuration through an exploratory study, e.g., increase of #category/#instance, patch rendering, image coloring, and training epoch. Although models pre-trained with the proposed Fractal DataBase (FractalDB), a database without natural images, do not necessarily outperform models pre-trained with human annotated datasets in all settings, we are able to partially surpass the accuracy of ImageNet/Places pre-trained models. The FractalDB pre-trained CNN also outperforms other pre-trained models on auto-generated datasets based on FDSL such as Bezier curves and Perlin noise. This is reasonable since natural objects and scenes existing around us are constructed according to fractal geometry. Image representation with the proposed FractalDB captures a unique feature in the visualization of convolutional layers and attentions. Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
Int. J. Comput. Vis. | 6 |
| 2021 | Teacher-Assisted Mini-Batch Sampling for Blind Distillation Using Metric LearningabstractThis paper addresses the problem of blind distillation, which aims to train a student model with unlabeled data under the supervision of a pre-trained teacher model. The proposed framework introduces metric learning to blind distillation. Specifically, teacher-assisted mini-batch (TAM) sampling is proposed, which makes triplets of anchor, positive and negative samples on unlabeled data by using the teacher's knowledge. In addition, we propose a metric-based loss, namely Contrastive Additive Margin (CAM) Softmax loss, which efficiently uses all combinations of triplets on each mini-batch obtained by TAM sampling. In experiments, we show the effectiveness of the proposed framework on face and speaker verification tasks, where student models are trained on unlabeled VoxCeleb videos with a teacher model pre-trained on VGGFace2 images. Nakamasa Inoue |
ICASSP | 1 |
| 2021 | Disentangling Latent Groups Of FactorsabstractThis paper proposes a framework for training variational autoencoders (VAEs) for image distributions that have latent groups of factors. Our key idea is to introduce a mechanism to predict the factor group an image belongs to while simultaneously disentangling factors in it. More specifically, we propose an architecture consisting of three components: an encoder, a decoder, and a factor-group prediction header. The first two components are trained with a VAE objective, and the last one is trained with the proposed algorithm using the loss of unsupervised contrastive learning. In experiments, we designed a task in which more than one group of factors were entangled by combining multiple datasets and demonstrated the effectiveness of the proposed framework. The Mutual Information Gap score was improved from 0.089 to 0.125 on a merged dataset of Color-dSprites, 3DShapes, and MPI3D. Nakamasa Inoue, Ryota Yamada, Rei Kawakami, Ikuro Sato |
ICIP | 1 |
| 2020 | Pre-training Without Natural Images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
ACCV (6) | 6 |
| 2020 | Initialization Using Perlin Noise for Training Networks with a Limited Amount of DataabstractWe propose a novel network initialization method using Perlin noise for training image classification networks with a limited amount of data. Our main idea is to initialize the network parameters by solving an artificial noise classification problem, where the aim is to classify Perlin noise samples into their noise categories. Specifically, the proposed method consists of two steps. First, it generates Perlin noise samples with category labels defined based on noise complexity. Second, it solves a classification problem, in which network parameters are optimized to classify the generated noise samples. This method produces a reasonable set of initial weights (filters) for image classification. To the best of our knowledge, this is the first work to initialize networks by solving an artificial optimization problem without using any real-world images. Our experiments show that the proposed method outperforms conventional initialization methods on four image classification datasets. Nakamasa Inoue, Eisuke Yamagata, Hirokatsu Kataoka |
ICPR | 1 |
| 2020 | Augmented Cyclic Consistency Regularization for Unpaired Image-to-Image TranslationabstractUnpaired image-to-image (I2I) translation has received considerable attention in pattern recognition and computer vision because of recent advancements in generative adversarial networks (GANs). However, due to the lack of explicit supervision, unpaired I2I models often fail to generate realistic images, especially in challenging datasets with different backgrounds and poses. Hence, stabilization is indispensable for GANs and applications of I2I translation. Herein, we propose Augmented Cyclic Consistency Regularization (ACCR), a novel regularization method for unpaired I2I translation. Our main idea is to enforce consistency regularization originating from semi-supervised learning on the discriminators leveraging real, fake, reconstructed, and augmented samples. We regularize the discriminators to output similar predictions when fed pairs of original and perturbed images. We qualitatively clarify why consistency regularization on fake and reconstructed samples works well. Quantitatively, our method outperforms the consistency regularized GAN (CR-GAN) in real-world translations and demonstrates efficacy against several data augmentation variants and cycle-consistent constraints. Takehiko Ohkawa, Naoto Inoue, Hirokatsu Kataoka, Nakamasa Inoue |
ICPR | 4 |
| 2020 | Graph Grouping Loss for Metric Learning of Face Image RepresentationsabstractThis paper proposes Graph Grouping (GG) loss for metric learning and its application to face verification. GG loss predisposes image embeddings of the same identity to be close to each other, and those of different identities to be far from each other by constructing and optimizing graphs representing the relation between images. Further, to reduce the computational cost, we propose an efficient way to compute GG loss for cases where embeddings are L2normalized. In experiments, we demonstrate the effectiveness of the proposed method for face verification on the VoxCeleb dataset. The results show that the proposed GG loss outperforms conventional losses for metric learning. Nakamasa Inoue |
VCIP | 1 |
| 2019 | Sequence-level Knowledge Distillation for Model Compression of Attention-based Sequence-to-sequence Speech RecognitionabstractWe investigate the feasibility of sequence-level knowledge distillation of Sequence-to-Sequence (Seq2Seq) models for Large Vocabulary Continuous Speech Recognition (LVCSR). We first use a pre-trained larger teacher model to generate multiple hypotheses per utterance with beam search. With the same input, we then train the student model using these hypotheses generated from the teacher as pseudo labels in place of the original ground truth labels. We evaluate our proposed method using Wall Street Journal (WSJ) corpus. It achieved up to 9.8× parameter reduction with accuracy loss of up to 7.0% word-error rate (WER) increase. Raden Muaz, Nakamasa Inoue, Koichi Shinoda |
ICASSP | 2 |
| 2018 | A Fine-to-Coarse Convolutional Neural Network for 3D Human Action Recognition
Thao Le Minh, Nakamasa Inoue, Koichi Shinoda |
BMVC | 2 |
| 2018 | Multi-Task Autoencoder for Noise-Robust Speech RecognitionabstractFor speech recognition in noisy environments, we propose a multi-task autoencoder which estimates not only clean speech features but also noise features from noisy speech. We introduce the deSpeeching autoencoder, which excludes speech signals from noisy speech, and combine it with the conventional denoising autoencoder to form a unified multi-task au-toencoder (MTAE). We evaluate it using the Aurora 2 dataset and CHIME 3 dataset. It reduced WER by 15.7% from the conventional denoising autoencoder in the Aurora 2 test set A. Haoyi Zhang, Conggui Liu, Nakamasa Inoue, Koichi Shinoda |
ICASSP | 3 |
| 2018 | Detecting Alzheimer's Disease Using Gated Convolutional Neural Network from Audio DataabstractWe propose an automatic detection method of Alzheimer's diseases using a gated convolutional neural network (GCNN) from speech data. This GCNN can be trained with a relatively small amount of data and can capture the temporal information in audio paralinguistic features. Since it does not utilize any linguistic features, it can be easily applied to any languages. We evaluated our method using Pitt Corpus. The proposed method achieved the accuracy of 73.6%, which is better than the conventional sequential minimal optimization (SMO) by 7.6 points. Tifani Warnita, Nakamasa Inoue, Koichi Shinoda |
INTERSPEECH | 2 |
| 2018 | I-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker VerificationabstractI-vector based text-independent speaker verification (SV) systems often have poor performance with short utterances, as the biased phonetic distribution in a short utterance makes the extracted i-vector unreliable.This paper proposes an i-vector compensation method using a generative adversarial network (GAN), where its generator network is trained to generate a compensated i-vector from a short-utterance i-vector and its discriminator network is trained to determine whether an i-vector is generated by the generator or the one extracted from a long utterance.Additionally, we assign two other learning tasks to the GAN to stabilize its training and to make the generated ivector more speaker-specific.Speaker verification experiments on the NIST SRE 2008 "10sec-10sec" condition show that after applying our method, the equal error rate reduced by 11.3% from the conventional i-vector and PLDA system. Jiacen Zhang, Nakamasa Inoue, Koichi Shinoda |
INTERSPEECH | 2 |
| 2018 | Few-Shot Adaptation for Multimedia Semantic IndexingabstractWe propose a few-shot adaptation framework, which bridges zero-shot learning and supervised many-shot learning, for semantic indexing of image and video data. Few-shot adaptation provides robust parameter estimation with few training examples, by optimizing the parameters of zero-shot learning and supervised many-shot learning simultaneously. In this method, first we build a zero-shot detector, and then update it by using the few examples. Our experiments show the effectiveness of the proposed framework on three datasets: TRECVID Semantic Indexing 2010, 2014, and ImageNET. On the ImageNET dataset, we show that our method outperforms recent few-shot learning methods. On the TRECVID 2014 dataset, we achieve 15.19~% and 35.98~% in Mean Average Precision under the zero-shot condition and the supervised condition, respectively. To the best of our knowledge, these are the best results on this dataset. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 1 |
| 2017 | Cross-view human action recognition from depth maps using spectral graph sequences
Tommi Kerola, Nakamasa Inoue, Koichi Shinoda |
Comput. Vis. Image Underst. | 2 |
| 2016 | Adaptation of Word Vectors using Tree Structure for Visual SemanticsabstractWe propose a framework of word-vector adaptation, which makes vectors of visually similar concepts close to each other. Here, word vectors are real-valued vector representation of words, e.g., word2vec representation. Our basic idea is to assume that each concept has some hypernyms that are important to determine its visual features. For example, for a concept Swallow with hypernyms Bird, Animal and Entity, we believe Bird is the most important since birds have common visual features with their feathers etc. Adapted word vectors are obtained for each word by taking a weighted sum of a given original word vector and its hypernym word vectors. Our weight optimization makes vectors of visually similar concepts close to each other, by giving a large weight for such important hypernyms. We apply the adapted word vectors to zero-shot learning on the TRECVID 2014 semantic indexing dataset. We achieved 0.083 of Mean Average Precision, which is the best performance without using TRECVID training data to the best of our knowledge. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 1 |
| 2016 | Fast Coding of Feature Vectors Using Neighbor-to-Neighbor SearchabstractSearching for matches to high-dimensional vectors using hard/soft vector quantization is the most computationally expensive part of various computer vision algorithms including the bag of visual word (BoW). This paper proposes a fast computation method, Neighbor-to-Neighbor (NTN) search [1] , which skips some calculations based on the similarity of input vectors. For example, in image classification using dense SIFT descriptors, the NTN search seeks similar descriptors from a point on a grid to an adjacent point. Applications of the NTN search to vector quantization, a Gaussian mixture model, sparse coding, and a kernel codebook for extracting image or video representation are presented in this paper. We evaluated the proposed method on image and video benchmarks: the PASCAL VOC 2007 Classification Challenge and the TRECVID 2010 Semantic Indexing Task. NTN-VQ reduced the coding cost by 77.4 percent, and NTN-GMM reduced it by 89.3 percent, without any significant degradation in classification performance. Nakamasa Inoue, Koichi Shinoda |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Vocabulary Expansion Using Word Vectors for Video Semantic IndexingabstractWe propose vocabulary expansion for video semantic indexing. From many semantic concept detectors obtained by using training data, we make detectors for concepts not included in training data. First, we introduce Mikolov's word vectors to represent a word by a low-dimensional vector. Second, we represent a new concept by a weighted sum of concepts in training data in the word vector space. Finally, we use the same weighting coefficients for combining detectors to make a new detector. In our experiments, we evaluate our methods on the TRECVID Video Semantic Indexing (SIN) Task. We train our models with Google News text documents and ImageNET images to generate new semantic detectors for SIN task. We show that our method performs as well as SVMs trained with 100 TRECVID ex- ample videos. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 1 |
| 2014 | Spectral Graph Skeletons for 3D Action Recognition
Tommi Kerola, Nakamasa Inoue, Koichi Shinoda |
ACCV (4) | 2 |
| 2014 | n-gram Models for Video Semantic IndexingabstractWe propose n-gram modeling of shot sequences for video semantic indexing, in which semantic concepts are extracted from a video shot. Most previous studies for this task have assumed that video shots in a video clip are independent from each other. We model the time-dependency between them assuming that n-consecutive video shots are dependent. Our models improve the robustness against occlusion and camera-angle changes by effectively using information from the previous video shots. In our experiments on the TRECVID 2012 Semantic Indexing Benchmark, we applied the proposed models to a system using Gaussian mixture models and support vector machines. Mean average precision was improved from 30.62% to 32.14%, which is the best performance on the TRECVID 2012 Semantic Indexing to the best of our knowledge. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 1 |
| 2014 | Event Detection by Velocity Pyramid
Zhuolin Liang, Nakamasa Inoue, Koichi Shinoda |
MMM (1) | 2 |
| 2013 | Neighbor-to-Neighbor Search for Fast Coding of Feature VectorsabstractAssigning a visual code to a low-level image descriptor, which we call code assignment, is the most computationally expensive part of image classification algorithms based on the bag of visual word (BoW) framework. This paper proposes a fast computation method, Neighbor-to-Neighbor (NTN) search, for this code assignment. Based on the fact that image features from an adjacent region are usually similar to each other, this algorithm effectively reduces the cost of calculating the distance between a codeword and a feature vector. This method can be applied not only to a hard codebook constructed by vector quantization (NTN-VQ), but also to a soft codebook, a Gaussian mixture model (NTN-GMM). We evaluated this method on the PASCAL VOC 2007 classification challenge task. NTN-VQ reduced the assignment cost by 77.4% in super-vector coding, and NTN-GMM reduced it by 89.3% in Fisher-vector coding, without any significant degradation in classification performance. Nakamasa Inoue, Koichi Shinoda |
ICCV | 1 |
| 2013 | q-Gaussian mixture models for image and video semantic indexing
Nakamasa Inoue, Koichi Shinoda |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | q-Gaussian Mixture Models Based on Non-extensive Statistics for Image and Video Semantic Indexing
Nakamasa Inoue, Koichi Shinoda |
ACCV (2) | 1 |
| 2012 | Multimedia event detection using GMM supervectors and SVMSabstractIn multimedia event detection, complex target events are extracted from a large set of consumer-generated videos taken in unconstrained environments. We devised a multimedia event detection method based on GMM supervectors and support vector machines (SVMs) using multiple features. A GMM supervector consists of the parameters of a Gaussian mixture model (GMM) for the distribution of local features extracted from a video clip. A GMM is regarded as an extension of the Bag-of-Words (BoW) to a probabilistic framework, and thus, it can be expected to be robust against the data insufficiency problem. This method outperformed previous methods including BoW in experiments using the dataset of the multimedia event detection task in TRECVID2010 and 2011. Yusuke Kamishima, Nakamasa Inoue, Koichi Shinoda, Shunsuke Sato |
ICIP | 2 |
| 2012 | A Fast and Accurate Video Semantic-Indexing System Using Fast MAP Adaptation and GMM SupervectorsabstractWe propose a fast maximum a posteriori (MAP) adaptation method for video semantic indexing that uses Gaussian mixture model (GMM) supervectors. In this method, a tree-structured GMM is utilzed to decrease the computational cost, where only the output probabilities of mixture components close to an input sample are precisely calculated. Experimental evaluation on the TRECVID 2010 dataset demonstrates the effectiveness of the proposed method. The calculation time of the MAP adaptation step is reduced by 76.2% compared with that of a conventional method. The total calculation time is reduced by 56.6% while keeping the same level of the accuracy. Nakamasa Inoue, Koichi Shinoda |
IEEE Trans. Multim. | 1 |
| 2011 | A fast MAP adaptation technique for gmm-supervector-based video semantic indexing systemsabstractWe propose a fast maximum a posteriori (MAP) adaptation technique for a GMM-supervectors-based video semantic indexing system.The use of GMM supervectors is one of the state-of-the-art methods in which MAP adaptation is needed for estimating the distribution of local features extracted from video data. The proposed method cuts the calculation time of the MAP adaptation step. With the proposed method, a tree-structured GMM is constructed to quickly calculate posterior probabilities for each mixture component of a GMM. The basic idea of the tree-structured GMM is to cluster Gaussian components and approximate them with a single Gaussian. Leaf nodes of the tree correspond to the mixture components, and each non-leaf node has a single Gaussian that approximates its descendant Gaussian distributions. Experimental evaluation on the TRECVID 2010 dataset demonstrates the effectiveness of the proposed method. The calculation time of the MAP adaptation step is reduced by 76.2% compared to that of a conventional method and resulting accuracy (in terms of Mean average precision) was 10.2%. Nakamasa Inoue, Koichi Shinoda |
ACM Multimedia | 1 |
| 2010 | High-Level Feature Extraction Using SIFT GMMs and Audio ModelsabstractWe propose a statistical framework for high-level feature extraction that uses SIFT Gaussian mixture models (GMMs) and audio models. SIFT features were extracted from all the image frames and modeled by a GMM. In addition, we used mel-frequency cepstral coefficients and ergodic hidden Markov models to detect high-level features in audio streams. The best result obtained by using SIFT GMMs in terms of mean average precision on the TRECVID 2009 corpus was 0.150 and was improved to 0.164 by using audio information. Nakamasa Inoue, Tatsuhiko Saito, Koichi Shinoda, Sadaoki Furui |
ICPR | 1 |