EDBT 2026 Demo / reviewers in the wild / expert
Rei Kawakami
dblp:69/6882
· DBLP profile ↗
48ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0003-2342-3324ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 25 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 15 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo Under Limited Multi-Illumination CuesabstractUniversal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unreliable, such as under biased lighting or in shadows or self-occluded regions of complex in-the-wild scenes. We propose GeoUniPS, a universal photometric stereo network that integrates synthetic supervision with high-level geometric priors from large-scale 3D reconstruction models pretrained on massive in-the-wild data. Our key insight is that these 3D reconstruction models serve as visual-geometry foundation models, inherently encoding rich geometric knowledge of real scenes. To leverage this, we design a Light-Geometry Dual-Branch Encoder that extracts both multi-illumination cues and geometric priors from the frozen 3D reconstruction model. We also address the limitations of the conventional orthographic projection assumption by introducing the PS-Perp dataset with realistic perspective projection to enable learning of spatially varying view directions. Extensive experiments demonstrate that GeoUniPS delivers state-of-the-arts performance across multiple datasets, both quantitatively and qualitatively, especially in the complex in-the-wild scenes. King-Man Tam, Satoshi Ikehata, Yuta Asano, Zhaoyi An, Rei Kawakami |
AAAI | 5 |
| 2025 | Rectified Lagrangian for Out-of-Distribution Detection in Modern Hopfield NetworksabstractModern Hopfield networks (MHNs) have recently gained significant attention in the field of artificial intelligence because they can store and retrieve a large set of patterns with an exponentially large memory capacity. A MHN is generally a dynamical system defined with Lagrangians of memory and feature neurons,where memories associated with in-distribution (ID) samples are represented by attractors in the feature space. One major problem in existing MHNs lies in managing out-of-distribution (OOD) samples because it was originally assumed that all samples are ID samples. To address this, we propose the rectified Lagrangian (RegLag), a new Lagrangian for memory neurons that explicitly incorporates an attractor for OOD samples in the dynamical system of MHNs. RecLag creates a trivial point attractor for any interaction matrix, enabling OOD detection by identifying samples that fall into this attractor as OOD. The interaction matrix is optimized so that the probability densities can be estimated to identify ID/OOD. We demonstrate the effectiveness of RecLag-based MHNs compared to energy-based OOD detection methods, including those using state-of-the-art Hopfield energies, across nine image datasets. Ryo Moriai, Nakamasa Inoue, Masayuki Tanaka 0001, Rei Kawakami, Satoshi Ikehata, Ikuro Sato |
AAAI | 4 |
| 2025 | Multi-Point Positional Insertion Tuning for Small Object DetectionabstractSmall object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensive. To address this issue, this paper introduces multi-point positional insertion (MPI) tuning, a parameter-efficient finetuning (PEFT) method for small object detection. Specifically, MPI incorporates multiple positional embeddings into a frozen pretrained model, enabling the efficient detection of small objects by providing precise positional information to latent features. Through experiments, we demonstrated the effectiveness of the proposed method on the SODA-D dataset. MPI performed comparably to conventional PEFT methods, including CoOp and VPT, while significantly reducing the number of parameters that need to be tuned. Kanoko Goto, Takumi Karasawa, Takumi Hirose, Rei Kawakami, Nakamasa Inoue |
ICASSP | 4 |
| 2025 | Binary Stochastic Flip Optimization for Training Binary Neural NetworksabstractFor deploying deep neural networks on edge devices with limited resources, binary neural networks (BNNs) have attracted significant attention, due to their computational and memory efficiency. However, once a neural network is binarized, finetuning it on edge devices becomes challenging because most conventional training algorithms for BNNs are designed for use on centralized servers and require storing real-valued parameters during training. To address this limitation, this paper introduces binary stochastic flip optimization (BinSFO), a novel training algorithm for BNNs. BinSFO employs a parameter update rule based on Boolean operations, eliminating the need to store real-valued parameters and thereby reducing memory requirements and computational overhead. In experiments, we demonstrated the effectiveness and memory efficiency of BinSFO in fine-tuning scenarios on six image classification datasets. BinSFO performed comparably to conventional training algorithms with a 70.7% smaller memory requirement. Code is released at https://github.com/TatsukichiShibuya/ICASSP2025_BinSFO Tatsukichi Shibuya, Nakamasa Inoue, Rei Kawakami, Ikuro Sato |
ICASSP | 3 |
| 2025 | Teach Me Sign: Stepwise Prompting LLM for Sign Language ProductionabstractLarge language models, with their strong reasoning ability and rich knowledge, have brought revolution to many tasks of AI, but their impact on sign language generation remains limited due to its complexity and unique rules. In this paper, we propose TEAch Me Sign (TEAM-Sign), treating sign language as another natural language. By fine-tuning an LLM, we enable it to learn the correspondence between text and sign language, and facilitate generation. Considering the differences between sign and spoken language, we employ a stepwise prompting strategy to extract the inherent sign language knowledge within the LLM, thereby supporting the learning and generation process. Experimental results on How2Sign and Phoenix14T datasets demonstrate that our approach effectively leverages both the sign language knowledge and reasoning capabilities of LLM to align the different distribution and grammatical rules between sign and spoken language. Zhaoyi An, Rei Kawakami |
ICIP | 2 |
| 2025 | A Unified Transformer-Based Framework with Pretraining for Whole Body Grasping Motion GenerationabstractWe present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications. The code can be seen$\color{Purple}{\text{here}}$. Edward Effendy, Kuan-Wei Tseng, Rei Kawakami |
ICIP | 3 |
| 2025 | Implicit Object Recognition via Reinforcement Learning in Out-Of-Domain ScenariosabstractObject detection in novel environments faces a significant challenge due to its reliance on extensive annotated datasets. This paper addresses this challenge by introducing a reinforcement learning (RL)-based framework that eliminates the need for labeled data while enabling implicit object detection. Our approach integrates an open-set object detector within the vision encoder of an RL agent, allowing detection to emerge naturally during task-driven interaction with the environment. By fine-tuning the vision model using Low-Rank Adaptation (LoRA), we achieve efficient and targeted adaptation to the unique visual characteristics of the environment without the computational burden of full retraining. Unlike traditional supervised methods, our RL-driven approach inherently learns to recognize and adapt to objects through interaction, achieving robust detection accuracy and frequency improvements for objects outside the scope of conventional pre-trained datasets. Experimental results demonstrate the efficacy of this annotation-free method, offering a scalable solution for object detection in diverse and evolving environments. Kenji Cari Koga, Rei Kawakami |
ICIP | 2 |
| 2025 | Collision Avoidance with Differentiable Occupancy Functions in Object RearrangementabstractWe address the challenge of object relocation by robots in environments where their behavior is expected to resemble that of humans. Existing methods typically learn to regress the position and orientation of objects specified by natural language commands using training data. However, these approaches do not account for physical constraints during training, often resulting in collisions between relocated objects. In this work, we introduce a collision avoidance loss based on functions that incorporate object size into the training process. Specifically, we propose a type of occupancy function in which particles are represented by a 3D Gaussian probability density function. By incorporating these functions into an additional training phase of existing models, we demonstrate a reduction in the number of collisions during rearrangement tasks. Notably, despite the decrease in collisions, the semantic structure of the relocation results is preserved. Roma Satoh, Nakamasa Inoue, Rei Kawakami |
IROS | 3 |
| 2025 | Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language ModelsabstractAchieving zero-shot peg insertion, where inserting an arbitrary peg into an unseen hole without task-specific training, remains a fundamental challenge in robotics. This task demands a highly generalizable perception system capable of detecting potential holes, selecting the correct mating hole from multiple candidates, estimating its precise pose, and executing insertion despite uncertainties. While learning-based methods have been applied to peg insertion, they often fail to generalize beyond the specific peg-hole pairs encountered during training. Recent advancements in Vision-Language Models (VLMs) offer a promising alternative, leveraging large-scale datasets to enable robust generalization across diverse tasks. Inspired by their success, we introduce a novel zero-shot peg insertion framework that utilizes a VLM to identify mating holes and estimate their poses without prior knowledge of their geometry. This approach assumes a known peg pose and a leveled surface for insertion. Extensive experiments demonstrate that our method achieves 90.2% accuracy, significantly outperforming baselines in identifying the correct mating hole across a wide range of previously unseen peg-hole pairs, including 3D-printed objects, toy puzzles, and industrial connectors. Furthermore, we validate the effectiveness of our approach in a real-world connector insertion task on a backpanel of a PC, where our system successfully detects holes, identifies the correct mating hole, estimates its pose, and completes the insertion with a success rate of 88.3%. These results highlight the potential of VLM-driven zero-shot reasoning for enabling robust and generalizable robotic assembly. Masaru Yajima, Kei Ota, Asako Kanezaki, Rei Kawakami |
IROS | 4 |
| 2025 | Contracted Gram Tensor Distillation for Object DetectionabstractKnowledge distillation (KD) is a technique for compressing large models into lightweight ones while preserving performance and has also proven effective in object detection. However, most existing detection KD methods naively adapt techniques designed for classification tasks, overlooking the uniqueness of the task output space in detection, including batch, class, and spatial dimensions. To address this limitation, we propose Contracted Gram Tensor Distillation (CGTD), a novel knowledge distillation framework for object detection. CGTD effectively distills structural knowledge in high-order tensors by minimizing loss between contracted Gram tensors, each of which is calculated along each axis of the target tensors and efficiently captures structural information. This enables the student model to learn not just the output values but also the underlying structural patterns, maximizing the effectiveness of knowledge distillation in object detection. We conduct experiments on the COCO and SODA-D datasets and show that our framework achieves state-of-the-art results in object detection distillation. Takumi Karasawa, Nakamasa Inoue, Rei Kawakami |
MMAsia | 3 |
| 2024 | Efficient Target Propagation by Deriving Analytical SolutionabstractExploring biologically plausible algorithms as alternatives to error backpropagation (BP) is a challenging research topic in artificial intelligence. It also provides insights into the brain's learning methods. Recently, when combined with well-designed feedback loss functions such as Local Difference Reconstruction Loss (LDRL) and through hierarchical training of feedback pathway synaptic weights, Target Propagation (TP) has achieved performance comparable to BP in image classification tasks. However, with an increase in the number of network layers, the tuning and training cost of feedback weights escalates. Drawing inspiration from the work of Ernoult et al., we propose a training method that seeks the optimal solution for feedback weights. This method enhances the efficiency of feedback training by analytically minimizing feedback loss, allowing the feedback layer to skip certain local training iterations. More specifically, we introduce the Jacobian matching loss (JML) for feedback training. We also proactively implement layers designed to derive analytical solutions that minimize JML. Through experiments, we have validated the effectiveness of this approach. Using the CIFAR-10 dataset, our method showcases accuracy levels comparable to state-of-the-art TP methods. Furthermore, we have explored its effectiveness in more intricate network architectures. Yanhao Bao, Tatsukichi Shibuya, Ikuro Sato, Rei Kawakami, Nakamasa Inoue |
AAAI | 4 |
| 2024 | Learning Non-uniform Step Sizes for Neural Network Quantization
Shinya Gongyo, Jinrong Liang, Mitsuru Ambai, Rei Kawakami, Ikuro Sato |
ACCV (8) | 4 |
| 2024 | A Simple Finetuning Strategy Based on Bias-Variance Ratios of Layer-Wise Gradients
Mao Tomita, Ikuro Sato, Rei Kawakami, Nakamasa Inoue, Satoshi Ikehata, Masayuki Tanaka 0001 |
ACCV (8) | 3 |
| 2024 | Cubic Knowledge Distillation for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) can play an important role in human-computer interaction. In this paper, we propose a logit knowledge distillation method for SER, called Cubic KD, that distill the knowledge of fine-tuned self-supervised models to allow better performance of small models. By creating cubic structures from teacher and student network output features and using a loss function to distill the cube structure through self-correlation between elements, Cubic KD efficiently captures knowledge within instances and among instances. We apply this distillation method to four student models and conduct experiments using the Emo-DB and IEMOCAP datasets. The results show that Cubic KD outperforms existing predictive logit knowledge distillation methods and is comparable to intermediate feature knowledge distillation methods. Our implementation code is available at https://github.com/Fly1toMoon/Cubic-Knowledge-Distillation Zhibo Lou, Shinta Otake, Zhengxiao Li, Rei Kawakami, Nakamasa Inoue |
ICASSP | 4 |
| 2024 | Spatiality-Aware Prompt Tuning for Few-Shot Small Object DetectionabstractSmall Object Detection (SOD) is challenging due to the scarcity of image features arising from small image regions. The niche nature of small objects additionally poses difficulty in data collection compared to normal-sized objects. Therefore, efficient learning from limited data is benefical for SOD. To tackle few-shot SOD, we propose Spatiality-Aware Prompt Tuning (SAPT), a novel prompt tuning method for vision-language models (VLMs) to deal with the image feature scarcity and the limited data for small objects. SAPT appends the verbalizer prompt, expressing the spatiality of small objects through a template-based sentence, to the text prompt of the pre-trained VLMs. During fine-tuning, the integrated text prompt is learned solely by the decoder of the vision-language detector, while the image and text backbones of the model remain frozen to facilitate efficient learning. In our experiments, we demonstrate the effectiveness of the proposed method on SODA-D and COCO datasets in few-shot and full-shot learning scenarios, and show that our method improves state-of-the-art in both scenarios. Takumi Karasawa, Nakamasa Inoue, Rei Kawakami |
ICIP | 3 |
| 2024 | Gumbel-NeRF: Representing Unseen Objects as Part-Compositional Neural Radiance FieldsabstractWe propose Gumbel-NeRF, a mixture-of-expert (MoE) neural radiance fields (NeRF) model with a hindsight expert selection mechanism for synthesizing novel views of unseen objects. Previous studies have shown that the MoE structure provides high-quality representations of a given large-scale scene consisting of many objects. However, we observe that such a MoE NeRF model often produces low-quality representations in the vicinity of experts’ boundaries when applied to the task of novel view synthesis of an unseen object from one/few-shot input. We find that this deterioration is primarily caused by the foresight expert selection mechanism, which may leave an unnatural discontinuity in the object shape near the experts’ boundaries. Gumbel-NeRF adopts a hindsight expert selection mechanism, which guarantees continuity in the density field even near the experts’ boundaries. Experiments using the SRN cars dataset demonstrate the superiority of Gumbel-NeRF over the baselines in terms of various image quality metrics. The code will be available upon acceptance. Yusuke Sekikawa, Chingwei Hsu, Satoshi Ikehata, Rei Kawakami, Ikuro Sato |
ICIP | 4 |
| 2024 | Object Detection Framework Using Multiple Tone Mappings on High-Dynamic-Range ImagesabstractIn practical computer vision applications, such as autonomous driving, the ability to effectively process high-dynamic-range (HDR) scenes is crucial for safe operation. In this paper, we focus on object detection within HDR images. To address this, we propose a simple yet effective framework that employs multiple tone mappings. First, we generate multiple images from an HDR image with varying tone mapping parameters. Then, those images are fed into a high-performance object detector pre-trained with low-dynamic-range (LDR) images. Multiple detection results are merged with non-maximum suppression (NMS). To assess the performance of our method, we have built a validation dataset comprising HDR images captured in outdoor scenes with significant contrast variations. The experimental results using both our dataset and an existing one demonstrate that our method outperforms existing approaches1.1The code and the dataset can be available at https://open-vision.sc.e.titech.ac.jp/research/hdrdet. Takumi Watanabe, Rei Kawakami, Masayuki Tanaka 0001, Masatoshi Okutomi |
ICIP | 2 |
| 2024 | Transferring Teacher's Invariance to Student Through Data Augmentation Optimization
Tamotsu Kurioka, Teppei Suzuki, Rei Kawakami, Ikuro Sato |
ICONIP (7) | 3 |
| 2024 | Few-Shot View Synthesis Based on Geometric and Semantic Consistency
Mizuki Kojima, Rei Kawakami, Masatoshi Okutomi |
ICPR (18) | 2 |
| 2024 | ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing TasksabstractSelf-supervised learning has emerged as a key approach for learning generic representations from speech data. Despite promising results in downstream tasks such as speech recognition, speaker verification, and emotion recognition, a significant number of parameters is required, which makes fine-tuning for each task memory-inefficient. To address this limitation, we introduce ELP-adapter tuning, a novel method for parameter-efficient fine-tuning using three types of adapter, namely encoder adapters (E-adapters), layer adapters (L-adapters), and a prompt adapter (P-adapter). The E-adapters are integrated into transformer-based encoder layers and help to learn finegrained speech representations that are effective for speech recognition. The L-adapters create paths from each encoder layer to the downstream head and help to extract non-linguistic features from lower encoder layers that are effective for speaker verification and emotion recognition. The P-adapter appends pseudo features to CNN features to further improve effectiveness and efficiency. With these adapters, models can be quickly adapted to various speech processing tasks. Our evaluation across four downstream tasks using five backbone models demonstrated the effectiveness of the proposed method. With the WavLM backbone, its performance was comparable to or better than that of full fine-tuning on all tasks while requiring 90% fewer learnable parameters. Nakamasa Inoue, Shinta Otake, Takumi Hirose, Masanari Ohi, Rei Kawakami |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Fixed-Weight Difference Target PropagationabstractTarget Propagation (TP) is a biologically more plausible algorithm than the error backpropagation (BP) to train deep networks, and improving practicality of TP is an open issue. TP methods require the feedforward and feedback networks to form layer-wise autoencoders for propagating the target values generated at the output layer. However, this causes certain drawbacks; e.g., careful hyperparameter tuning is required to synchronize the feedforward and feedback training, and frequent updates of the feedback path are usually required than that of the feedforward path. Learning of the feedforward and feedback networks is sufficient to make TP methods capable of training, but is having these layer-wise autoencoders a necessary condition for TP to work? We answer this question by presenting Fixed-Weight Difference Target Propagation (FW-DTP) that keeps the feedback weights constant during training. We confirmed that this simple method, which naturally resolves the abovementioned problems of TP, can still deliver informative target values to hidden layers for a given task; indeed, FW-DTP consistently achieves higher test performance than a baseline, the Difference Target Propagation (DTP), on four classification datasets. We also present a novel propagation architecture that explains the exact form of the feedback function of DTP to analyze FW-DTP. Our code is available at https://github.com/TatsukichiShibuya/Fixed-Weight-Difference-Target-Propagation. Tatsukichi Shibuya, Nakamasa Inoue, Rei Kawakami, Ikuro Sato |
AAAI | 3 |
| 2023 | Learning with Partial Forgetting in Modern Hopfield NetworksabstractIt has been known by neuroscience studies that partial and transient forgetting of memory often plays an important role in the brain to improve performance for certain intellectual activities. In machine learning, associative memory models such as classical and modern Hopfield networks have been proposed to express memories as attractors in the feature space of a closed recurrent network. In this work, we propose learning with partial forgetting (LwPF), where a partial forgetting functionality is designed by element-wise non-bijective projections, for memory neurons in modern Hopfield networks to improve model performance. We incorporate LwPF into the attention mechanism also, whose process has been shown to be identical to the update rule of a certain modern Hopfield network, by modifying the corresponding Lagrangian. We evaluated the effectiveness of LwPF on three diverse tasks such as bit-pattern classification, immune repertoire classification for computational biology, and image classification for computer vision, and confirmed that LwPF consistently improves the performance of existing neural networks including DeepRC and vision transformers. Toshihiro Ota, Ikuro Sato, Rei Kawakami, Masayuki Tanaka 0001, Nakamasa Inoue |
AISTATS | 3 |
| 2023 | Step restriction for improving adversarial attacksabstractWe propose an algorithm to automatically restrict the step size in the iterative optimization process with an application to adversarial attacks on speaker verification models. The proposed algorithm dynamically determines a subspace with a restriction radius r to which the Taylor approximation is applied at each iteration and then solves a linear problem within the subspace by using the projected gradient method. In experiments, we demonstrate adversarial attacks on three speaker verification models: i-vectors, SE-ResNet-34, and ECAPATDNN. We show that the degree of adversarial perturbations generated by the proposed algorithm is smaller than that generated by the conventional attack method. Keita Goto, Shinta Otake, Rei Kawakami, Nakamasa Inoue |
ICASSP | 3 |
| 2023 | Parameter Efficient Transfer Learning for Various Speech Processing TasksabstractFine-tuning of self-supervised models is a powerful transfer learning method in a variety of fields, including speech processing, since it can utilize generic feature representations obtained from large amounts of unlabeled data. Fine-tuning, however, requires a new parameter set for each downstream task, which is parameter inefficient. Adapter architecture is proposed to partially solve this issue by inserting lightweight learnable modules into a frozen pre-trained model. However, existing adapter architectures fail to adaptively leverage low-to high-level features stored in different layers, which is necessary for solving various kinds of speech processing tasks. Thus, we propose a new adapter architecture to acquire feature representations more flexibly for various speech tasks. In experiments, we applied this adapter to WavLM on four speech tasks. It performed on par or better than naïve fine-tuning, with only 11% of learnable parameters. It also outperformed an existing adapter architecture. Our implementation code is available at https://github.com/sinhat98/adapter-wavlm Shinta Otake, Rei Kawakami, Nakamasa Inoue |
ICASSP | 2 |
| 2023 | Scale-space Tokenization for Improving the Robustness of Vision TransformersabstractThe performance of the Vision Transformer (ViT) model and its variants in most vision tasks has surpassed traditional Convolutional Neural Networks (CNNs) in terms of in-distribution accuracy. However, ViTs still have significant room for improvement in their robustness to input perturbations. Furthermore, robustness is a critical aspect to consider when deploying ViTs in real-world scenarios. Despite this, some variants of ViT improve the in-distribution accuracy and computation performance at the cost of sacrificing the model's robustness and generalization. In this study, inspired by the prior findings on the potential effectiveness of shape bias to robustness improvement and the importance of multi-scale analysis, we propose a simple yet effective method, scale-space tokenization, to improve the robustness of ViT while maintaining in-distribution accuracy. Based on this method, we build Scale-space-based Robust Vision Transformer (SRVT) model. Our method consists of scale-space patch embedding and scale-space positional encoding. The scale-space patch embedding makes a sequence of variable-scale images and increases the model's shape bias to enhance its robustness. The scale-space positional encoding implicitly boosts the model's invariance to input perturbations by incorporating scale-aware position information into 3D sinusoidal positional encoding. We conduct experiments on image recognition benchmarks (CIFAR10/100 and ImageNet-1k) from the perspectives of in-distribution accuracy, adversarial and out-of-distribution robustness. The experimental results demonstrate our method's effectiveness in improving robustness without compromising in-distribution accuracy. Especially, our approach achieves advanced adversarial robustness on ImageNet-1k benchmark compared with state-of-the-art robust ViT. Rei Kawakami, Nakamasa Inoue |
ACM Multimedia | 2 |
| 2023 | EvIs-Kitchen: Egocentric Human Activities Recognition with Video and Inertial Sensor Data
Yuzhe Hao, Kuniaki Uto, Asako Kanezaki, Ikuro Sato, Rei Kawakami, Koichi Shinoda |
MMM (1) | 5 |
| 2023 | Local Brightness Normalization for Image Classification and Object Detection Robust to Illumination Changes
Yanshuo Lu, Masayuki Tanaka 0001, Rei Kawakami, Masatoshi Okutomi |
PSIVT | 3 |
| 2022 | Multi-task Curriculum Learning based on Gradient Similarity
Hiroaki Igarashi, Kenichi Yoneji, Kohta Ishikawa, Rei Kawakami, Teppei Suzuki, Shingo Yashima, Ikuro Sato |
BMVC | 4 |
| 2022 | PoF: Post-Training of Feature Extractor for Improving GeneralizationabstractIt has been intensively investigated that the local shape, especially flatness, of the loss landscape near a minimum plays an important role for generalization of deep models. We developed a training algorithm called PoF: Post-Training of Feature Extractor that updates the feature extractor part of an already-trained deep model to search a flatter minimum. The characteristics are two-fold: 1) Feature extractor is trained under parameter perturbations in the higher-layer parameter space, based on observations that suggest flattening higher-layer parameter space, and 2) the perturbation range is determined in a data-driven manner aiming to reduce a part of test loss caused by the positive loss curvature. We provide a theoretical analysis that shows the proposed algorithm implicitly reduces the target Hessian components as well as the loss. Experimental results show that PoF improved model performance against baseline methods on both CIFAR-10 and CIFAR-100 datasets for only 10-epoch post-training, and on SVHN dataset for 50-epoch post-training. Ikuro Sato, Ryota Yamada, Masayuki Tanaka 0001, Nakamasa Inoue, Rei Kawakami |
ICML | 5 |
| 2022 | Feature Space Particle Inference for Neural Network EnsemblesabstractEnsembles of deep neural networks demonstrate improved performance over single models. For enhancing the diversity of ensemble members while keeping their performance, particle-based inference methods offer a promising approach from a Bayesian perspective. However, the best way to apply these methods to neural networks is still unclear: seeking samples from the weight-space posterior suffers from inefficiency due to the over-parameterization issues, while seeking samples directly from the function-space posterior often leads to serious underfitting. In this study, we propose to optimize particles in the feature space where activations of a specific intermediate layer lie to alleviate the abovementioned difficulties. Our method encourages each member to capture distinct features, which are expected to increase the robustness of the ensemble prediction. Extensive evaluation on real-world datasets exhibits that our model significantly outperforms the gold-standard Deep Ensembles on various metrics, including accuracy, calibration, and robustness. Shingo Yashima, Teppei Suzuki, Kohta Ishikawa, Ikuro Sato, Rei Kawakami |
ICML | 5 |
| 2022 | Informative Sample-Aware Proxy for Deep Metric LearningabstractAmong various supervised deep metric learning methods proxy-based approaches have achieved high retrieval accuracies. Proxies, which are class-representative points in an embedding space, receive updates based on proxy-sample similarities in a similar manner to sample representations. In existing methods, a relatively small number of samples can produce large gradient magnitudes (i.e., hard samples), and a relatively large number of samples can produce small gradient magnitudes (i.e., easy samples); these can play a major part in updates. Assuming that acquiring too much sensitivity to such extreme sets of samples would deteriorate the generalizability of a method, we propose a novel proxy-based method called Informative Sample-Aware Proxy (Proxy-ISA), which directly modifies a gradient weighting factor for each sample using a scheduled threshold function, so that the model is more sensitive to the informative samples. Extensive experiments on the CUB-200-2011, Cars-196, Stanford Online Products and In-shop Clothes Retrieval datasets demonstrate the superiority of Proxy-ISA compared with the state-of-the-art methods. Aoyu Li, Ikuro Sato, Kohta Ishikawa, Rei Kawakami, Rio Yokota |
MMAsia | 4 |
| 2021 | Disentangling Latent Groups Of FactorsabstractThis paper proposes a framework for training variational autoencoders (VAEs) for image distributions that have latent groups of factors. Our key idea is to introduce a mechanism to predict the factor group an image belongs to while simultaneously disentangling factors in it. More specifically, we propose an architecture consisting of three components: an encoder, a decoder, and a factor-group prediction header. The first two components are trained with a VAE objective, and the last one is trained with the proposed algorithm using the loss of unsupervised contrastive learning. In experiments, we designed a task in which more than one group of factors were entangled by combining multiple datasets and demonstrated the effectiveness of the proposed framework. The Mutual Information Gap score was improved from 0.089 to 0.125 on a merged dataset of Color-dSprites, 3DShapes, and MPI3D. Nakamasa Inoue, Ryota Yamada, Rei Kawakami, Ikuro Sato |
ICIP | 3 |
| 2020 | RNN-based Motion Prediction in Competitive Fencing Considering Interaction between Players
Yutaro Honda, Rei Kawakami, Takeshi Naemura |
BMVC | 2 |
| 2019 | Classification-Reconstruction Learning for Open-Set RecognitionabstractOpen-set classification is a problem of handling `unknown' classes that are not contained in the training dataset, whereas traditional classifiers assume that only known classes appear in the test environment. Existing open-set classifiers rely on deep networks trained in a supervised manner on known classes in the training set; this causes specialization of learned representations to known classes and makes it hard to distinguish unknowns from knowns. In contrast, we train networks for joint classification and reconstruction of input data. This enhances the learned representation so as to preserve information useful for separating unknowns from knowns, as well as to discriminate classes of knowns. Our novel Classification-Reconstruction learning for Open-Set Recognition (CROSR) utilizes latent representations for reconstruction and enables robust unknown detection without harming the known-class classification accuracy. Extensive experiments reveal that the proposed method outperforms existing deep open-set classifiers in multiple standard datasets and is robust to diverse outliers. Ryota Yoshihashi, Wen Shao, Rei Kawakami, Shaodi You, Makoto Iida, Takeshi Naemura |
CVPR | 3 |
| 2019 | Cross-Connected Networks for Multi-Task Learning of Detection and SegmentationabstractMulti-task learning improves generalization performance in neural networks by sharing knowledge among related tasks. Existing models are for task combinations annotated on the same dataset; research on how to utilize the knowledge of successful single-task convolutional neural networks (CNNs) that are trained on individual datasets is limited. We propose a cross-connected CNN, an architecture that connects single-task CNNs through convolutional layers that transfer useful information to their counterparts. We evaluated the architecture with a combination of detection and segmentation using datasets of two targets: pedestrians and wild birds. Experiments demonstrate how well our CNN learns general representations from multi-task learning. Rei Kawakami, Ryota Yoshihashi, Seiichiro Fukuda, Shaodi You, Makoto Iida, Takeshi Naemura |
ICIP | 1 |
| 2016 | Detection of small birds in large images by combining a deep detector with semantic segmentationabstractThis paper tackles the problem of bird detection in large landscape images for applications in the wind energy industry. While significant progress in image recognition has been made by deep convolutional neural networks (CNNs), small object detection remains a problem. To solve it, we follow the idea that a detector can be tuned to small objects of interest and semantic segmentation methods can be complementary used to recognize large background areas. Specifically, we train a CNN-based detector, fully convolutional networks, and a superpixel-based semantic segmentation method. The results of the three methods are combined by using support vector machines to achieve high detection performance. Experimental results on a bird image dataset show the high precision and effectiveness of the proposed method. Akito Takeki, Tu Tuan Trinh, Ryota Yoshihashi, Rei Kawakami, Makoto Iida, Takeshi Naemura |
ICIP | 4 |
| 2016 | Adherent Raindrop Modeling, Detectionand Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Modeling, detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. To accomplish the idea, we first model adherent raindrops using law of physics, and detect raindrops based on these models in combination with motion and intensity temporal derivatives of the input video. Having detected the raindrops, we remove them and restore the images based on an analysis that some areas of raindrops completely occludes the scene, and some other areas occlude only partially. For partially occluding areas, we restore them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity derivative. For completely occluding areas, we recover them by using a video completion technique. Experimental results using various real videos show the effectiveness of our method. Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Construction of a bird image dataset for ecological investigationsabstractWind turbines have become significant collision risks to endangered birds. To estimate the impact of bird strikes and take action against them, we need an automated bird monitoring system that detects and categorizes bird species. Towards this goal, we constructed an image dataset as a benchmark for ecological investigations of birds. The dataset consists of time-series images of a scene in a wind farm, bounding boxes around the aerial objects in the images, and species annotations. Bird experts searched for and annotated the images of birds, and thus, even birds that appeared to be very small in the whole image could be specified in detail. In total, 32,000 birds are annotated. In collecting the images, we had to monitor birds in the distance to cover enough area, but even with a telephoto setup, the average size of the birds in the image is around 25 pixels. This means we need a means of recognition that works on very low-resolution images. As an application of the dataset, we conducted experiments in two-class categorization (birds and non-birds), to serve as basis for detection and further categorization. We found that the dataset is a challenging recognition task because of its low resolution and hard negative samples. Ryota Yoshihashi, Rei Kawakami, Makoto Iida, Takeshi Naemura |
ICIP | 2 |
| 2014 | Raindrop Detection and Removal from Long Range Trajectories
Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
ACCV (2) | 3 |
| 2013 | Adherent Raindrop Detection and Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. First, it detects raindrops based on the motion and the intensity temporal derivatives of the input video. Second, relying on an analysis that some areas of a raindrop completely occludes the scene, yet the remaining areas occludes only partially, the method removes the two types of areas separately. For partially occluding areas, it restores them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity change. For completely occluding areas, it recovers them by using a video completion technique. Experimental results using various real videos show the effectiveness of the proposed method. Shaodi You, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi |
CVPR | 3 |
| 2013 | Camera Spectral Sensitivity and White Balance Estimation from Sky ImagesabstractPhotometric camera calibration is often required in physics-based computer vision. There have been a number of studies to estimate camera response functions (gamma function), and vignetting effect from images. However less attention has been paid to camera spectral sensitivities and white balance settings. This is unfortunate, since those two properties significantly affect image colors. Motivated by this, a method to estimate camera spectral sensitivities and white balance setting jointly from images with sky regions is introduced. The basic idea is to use the sky regions to infer the sky spectra. Given sky images as the input and assuming the sun direction with respect to the camera viewing direction can be extracted, the proposed method estimates the turbidity of the sky by fitting the image intensities to a sky model. Subsequently, it calculates the sky spectra from the estimated turbidity. Having the sky $$RGB$$ values and their corresponding spectra, the method estimates the camera spectral sensitivities together with the white balance setting. Precomputed basis functions of camera spectral sensitivities are used in the method for robust estimation. The whole method is novel and practical since, unlike existing methods, it uses sky images without additional hardware, assuming the geolocation of the captured sky is known. Experimental results using various real images show the effectiveness of the method. Rei Kawakami, Hongxun Zhao, Robby T. Tan, Katsushi Ikeuchi |
Int. J. Comput. Vis. | 1 |
| 2011 | High-resolution hyperspectral imaging via matrix factorizationabstractHyperspectral imaging is a promising tool for applications in geosensing, cultural heritage and beyond. However, compared to current RGB cameras, existing hyperspectral cameras are severely limited in spatial resolution. In this paper, we introduce a simple new technique for reconstructing a very high-resolution hyperspectral image from two readily obtained measurements: A lower-resolution hyper-spectral image and a high-resolution RGB image. Our approach is divided into two stages: We first apply an unmixing algorithm to the hyperspectral input, to estimate a basis representing reflectance spectra. We then use this representation in conjunction with the RGB input to produce the desired result. Our approach to unmixing is motivated by the spatial sparsity of the hyperspectral input, and casts the unmixing problem as the search for a factorization of the input into a basis and a set of maximally sparse coefficients. Experiments show that this simple approach performs reasonably well on both simulations and real data examples. Rei Kawakami, Yasuyuki Matsushita, John Wright 0001, Moshe Ben-Ezra, Yu-Wing Tai, Katsushi Ikeuchi |
CVPR | 1 |
| 2010 | Surface color estimation based on inter- and intra-pixel relationships in outdoor scenesabstractWe propose a method for estimating inherent surface color robustly against image noises from two registered images taken under different outdoor illuminations. We formulate the estimation based on maximum likelihood manner while considering both inter-pixel and intra-pixel relationships. We define inter-pixel relationship based on stochastic behavior of image noises and properties of outdoor illumination chromaticity. We rely on the spatial continuity of both surface color and illumination to define intra-pixel relationship. We also propose to maximize the estimation function in two step manner. Experimental results demonstrate the significant improvement of the proposed method in estimation accuracy compared to previous methods. Shun Hirose, Tsuyoshi Suenaga, Kentaro Takemura, Rei Kawakami, Jun Takamatsu, Tsukasa Ogasawara |
CVPR | 4 |
| 2010 | Estimating optical properties of layered surfaces using the spider modelabstractMany object surfaces are composed of layers of different physical substances, known as layered surfaces. These surfaces, such as patinas, water colors, and wall paintings, have more complex optical properties than diffuse surfaces. Although the characteristics of layered surfaces, like layer opacity, mixture of colors, and color gradations, are significant, they are usually ignored in the analysis of many methods in computer vision, causing inaccurate or even erroneous results. Therefore, the main goals of this paper are twofold: to solve problems of layered surfaces by focusing mainly on surfaces with two layers (i.e., top and bottom layers), and to introduce a decomposition method based on a novel representation of a nonlinear correlation in the color space that we call the “spider” model. When we plot a mixture of colors of one bottom layer and n different top layers into the RGB color space, then we will have n different curves intersecting at one point, resembling the shape of a spider. Hence, given a single input image containing one bottom layer and at least one top layer, we can fit their color distributions by using the spider model and then decompose those layered surfaces. The last step is equivalent to extracting the approximated optical properties of the two layers: the top layer's opacity, and the top and bottom layers' reflections. Experiments with real images, which include the photographs of ancient wall paintings, show the effectiveness of our method. Tetsuro Morimoto, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi |
CVPR | 3 |
| 2010 | Foreground and shadow occlusion handling for outdoor augmented realityabstractOcclusion handling in augmented reality (AR) applications is challenging in synthesizing virtual objects correctly into the real scene with respect to existing foregrounds and shadows. Furthermore, outdoor environment makes the task more difficult due to the unpredictable illumination changes. This paper proposes novel outdoor illumination constraints for resolving the foreground occlusion problem in outdoor environment. The constraints can be also integrated into a probabilistic model of multiple cues for a better segmentation of the foreground. In addition, we introduce an effective method to resolve the shadow occlusion problem by using shadow detection and recasting with a spherical vision camera. We have applied the system in our digital cultural heritage project named Virtual Asuka (VA) and verified the effectiveness of the system. Boun Vinh Lu, Tetsuya Kakuta, Rei Kawakami, Takeshi Oishi, Katsushi Ikeuchi |
ISMAR | 3 |
| 2009 | Color estimation from a single surface colorabstractThis paper estimates illumination colors by using only a single surface color taken under multiple illumination colors. Past researchers have found that there is a difficulty in estimating illumination colors using a single surface color. However, the method presented here overcomes the problem. Surface color is estimated by considering four characteristics of illumination and surface color spaces. First, the outdoor-illumination colors exist in a specific color range. Second, multiple illuminations give constraints for the surface color. Third, multiple illuminations also give constraints for the color range. Fourth, each color component affects those constraints in a different manner. Based on those characteristics, a novel method can be designed. The proposed method produces consistently accurate results when multiple illumination colors are used, because the constraint (possible range of illumination colors) on illumination colors refines the estimated illumination colors, effectively. Rei Kawakami, Katsushi Ikeuchi |
CVPR | 1 |
| 2008 | Detection of moving objects and cast shadows using a spherical vision camera for outdoor mixed realityabstractThis paper presents a method to detect moving objects and remove their shadows for superimposing them on Mixed Reality (MR) systems. We cut out the foreground from a real image using a probability-based segmentation method. Using color, spatial, and temporal priors, we can improve the accuracy of the segmentation. Energy minimization is executed by graph cuts. Then we remove the shadow region from the foreground with F-value calculated from the pixel value and the spectral sensitivity characteristic of the camera. Finally we superimpose virtual objects using the stencil buffer, which is used to limit the area of rendering for each pixel. Synthesized images of an outdoor scene show the efficiency of the proposed method. Tetsuya Kakuta, Boun Vinh Lu, Rei Kawakami, Takeshi Oishi, Katsushi Ikeuchi |
VRST | 3 |
| 2005 | Consistent Surface Color for Texturing Large Objects in Outdoor ScenesabstractColor appearance of an object is significantly influenced by the color of the illumination. When the illumination color changes, the color appearance of the object change accordingly, causing its appearance to be inconsistent. To arrive at color constancy, we have developed a physics-based method of estimating and removing the illumination color. In this paper, we focus on the use of this method to deal with outdoor scenes, since very few physics-based methods have successfully handled outdoor color constancy. Our method is principally based on shadowed and non-shadowed regions. Previously researchers have discovered that shadowed regions are illuminated by sky light, while non-shadowed regions are illuminated by a combination of sky light and sunlight. Based on this difference of illumination, we estimate the illumination colors (both the sunlight and the sky light) and then remove them. To reliably estimate the illumination colors in outdoor scenes, we include the analysis of noise, since the presence of noise is inevitable in natural images. As a result, compared to existing methods, the proposed method is more effective and robust in handling outdoor scenes. In addition, the proposed method requires only a single input image, making it useful for many applications of computer vision Rei Kawakami, Katsushi Ikeuchi, Robby T. Tan |
ICCV | 1 |