EDBT 2026 Demo / reviewers in the wild / expert
Shoichiro Takeda
dblp:230/1270
· DBLP profile ↗
16ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-7138-4983ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsabstractContrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics. Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura |
AAAI | 3 |
| 2026 | IPCD: Intrinsic Point-Cloud DecompositionabstractPoint clouds are widely used in various fields, including augmented reality (AR) and robotics, where relighting and texture editing are crucial for realistic visualization. Achieving these tasks requires accurately separating albedo from shade. However, performing this separation on point clouds presents two key challenges: (1) the non-grid structure of point clouds makes conventional image-based decomposition models ineffective, and (2) point-cloud models designed for other tasks do not explicitly consider global-light direction, resulting in inaccurate shade. In this paper, we introduce Intrinsic Point-Cloud Decomposition (IPCD), which extends image decomposition to the direct decomposition of colored point clouds into albedo and shade. To overcome challenge (1), we propose IPCD-Net that extends image-based model with point-wise feature aggregation for non-grid data processing. For challenge (2), we introduce Projection-based Luminance Distribution (PLD) with a hierarchical feature refinement, capturing global-light ques via multi-view projection. For comprehensive evaluation, we create a synthetic outdoor-scene dataset. Experimental results demonstrate that IPCD-Net reduces cast shadows in albedo and enhances color accuracy in shade. Furthermore, we showcase its applications in texture editing, relighting, and point-cloud registration under varying illumination. Finally, we verify the real-world applicability of IPCD-Net. Shogo Sato, Takuhiro Kaneko, Shoichiro Takeda, Tomoyasu Shimada, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida, Akisato Kimura |
WACV | 3 |
| 2026 | Distribution Highlighted Reference-based Label Distribution Learning for Facial Age EstimationabstractEstimating age from facial images is a fundamental task. In this task, age labels have ambiguity because faces of the same individual across similar ages are often difficult to distinguish. To model this ambiguity, label distribution learning (LDL) trains a deep neural network (DNN) using a label distribution, i.e., the probability that an image belongs to each age, instead of a single age label. However, the heuristic constraints utilized for LDL often fail to accurately model the label ambiguity. Therefore, we propose a novel LDL method called distribution highlighted reference-based LDL (DHRL), which introduces an input-dependent constraint by utilizing a reference DNN pre-trained with any LDL method and minimizing the gap between the reference and target DNNs’ outputs. DHRL incorporates two techniques to highlight the label ambiguity hidden in the pre-trained reference DNN’s output: noisy augmentation-based ensembling (NAE) and different scale multi-temperature (DSM). NAE inputs noisy images to the reference DNN and provides an ensemble effect by averaging all the outputs. DSM sets multiple temperatures simultaneously in the gap minimization between the two DNNs’ outputs. Experimental results indicate that our method achieves state-of-the-art performance across various datasets and conditions. Shin'ya Yamaguchi, Shoichiro Takeda, Takuhiro Kaneko, Shota Orihashi, Ryo Masumura |
WACV | 3 |
| 2025 | Improving Putting Accuracy with Electrical Muscle Stimulation Feedback Guided by Muscle Synergy Analysis
Arinobu Niijima, Shoichiro Takeda |
CHI | 2 |
| 2025 | Gromov-Wasserstein Problem with Cyclic SymmetryabstractWe propose novel fast algorithms for the Gromov–Wasserstein problem (GW) with cyclic symmetry of input data. This problem naturally appears as an object-matching task, which underlies various real-world computer vision applications, e.g., image registration, point cloud registration, stereo matching, and 3D reconstruction. Gradient-based algorithms have been widely used to solve GW, and our main idea is to utilize the following remarkable property that emerges in GW with cyclic symmetry: By setting the initial solution to have cyclic symmetry, all intermediate solutions and matrices that appear in the gradient-based algorithms have the same cyclic symmetry until convergence. Based on this property, our gradient-based algorithms restrict the solution space to have cyclic symmetry and update only one symmetric part of solutions and matrices at each iteration, resulting in faster computation. Moreover, our algorithms solve the optimal transport problem at each iteration, which also exhibits cyclic symmetry. This problem can be solved efficiently, and as a result, our algorithms perform significantly faster. Experiments showed the effectiveness of our algorithms in synthetic and real-world data with strict and approximate cyclic symmetry. Shoichiro Takeda, Yasunori Akagi |
CVPR | 1 |
| 2024 | Optimal Transport with Cyclic SymmetryabstractWe propose novel fast algorithms for optimal transport (OT) utilizing a cyclic symmetry structure of input data. Such OT with cyclic symmetry appears universally in various real-world examples: image processing, urban planning, and graph processing. Our main idea is to reduce OT to a small optimization problem that has significantly fewer variables by utilizing cyclic symmetry and various optimization techniques. On the basis of this reduction, our algorithms solve the small optimization problem instead of the original OT. As a result, our algorithms obtain the optimal solution and the objective function value of the original OT faster than solving the original OT directly. In this paper, our focus is on two crucial OT formulations: the linear programming OT (LOT) and the strongly convex-regularized OT, which includes the well-known entropy-regularized OT (EROT). Experiments show the effectiveness of our algorithms for LOT and EROT in synthetic/real-world data that has a strict/approximate cyclic symmetry structure. Through theoretical and experimental results, this paper successfully introduces the concept of symmetry into the OT research field for the first time. Shoichiro Takeda, Yasunori Akagi, Naoki Marumo, Kenta Niwa |
AAAI | 1 |
| 2023 | Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffabstractThis paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do. Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura |
ICCV | 3 |
| 2023 | Distorted image classification using neural activation pattern matching lossabstractIn image classification, a deep neural network (DNN) that is trained on undistorted images constitutes an effective decision boundary. Unfortunately, this boundary does not support distorted images, such as noisy or blurry ones, leading to accuracy drop-off. As a simple approach for classifying distorted images as well as undistorted ones, previous methods have optimized the trained DNN again on both kinds of images. However, in these methods, the decision boundary may become overly complicated during optimization because there is no regularization of the decision boundary. Consequently, this decision boundary limits efficient optimization. In this paper, we study a simple yet effective decision boundary for distorted image classification through the use of a novel loss, called a "neural activation pattern matching (NAPM) loss". The NAPM loss is based on recent findings that the decision boundary is a piecewise linear function, where each linear segment is constructed from a neural activation pattern in the DNN when an image is fed to it. The NAPM loss extracts the neural activation patterns when the distorted image and its undistorted version are fed to the DNN and then matches them with each other via the sigmoid cross-entropy. Therefore, it constrains the DNN to classify the distorted image and its undistorted version by the same linear segment. As a result, our loss accelerates efficient optimization by preventing the decision boundary from becoming overly complicated. Our experiments demonstrate that our loss increases the accuracy of the previous methods in all conditions evaluated. Shoichiro Takeda, Ryuichi Tanida, Yukihiro Bandoh, Hayaru Shouno |
Neural Networks | 2 |
| 2022 | Bilateral Video Magnification FilterabstractEulerian video magnification (EVM) has progressed to magnify subtle motions with a target frequency even under the presence of large motions of objects. However, existing EVM methods often fail to produce desirable results in real videos due to (1) misextracting subtle motions with a non-target frequency and (2) collapsing results when large de/acceleration motions occur (e.g., objects suddenly start, stop, or change direction). To enhance EVM performance on real videos, this paper proposes a bilateral video magnification filter (BVMF) that offers simple yet robust temporal filtering. BVMF has two kernels; (I) one kernel performs temporal bandpass filtering via a Laplacian of Gaussian whose passband peaks at the target frequency with unity gain and (II) the other kernel excludes large motions outside the magnitude of interest by Gaussian filtering on the intensity of the input signal via the Fourier shift theorem. Thus, BVMF extracts only subtle motions with the target frequency while excluding large motions outside the magnitude of interest, regardless of motion dynamics. In addition, BVMF runs the two kernels in the temporal and intensity domains simultaneously like the bilateral filter does in the spatial and intensity domains. This simplifies implementation and, as a secondary effect, keeps the memory usage low. Experiments conducted on synthetic and real videos show that BVMF outperforms state-of-the-art methods. Shoichiro Takeda, Kenta Niwa, Mariko Isogawa, Shinya Shimizu, Kazuki Okami, Yushi Aono |
CVPR | 1 |
| 2022 | Deep Feature Compression Using Spatio-Temporal Arrangement Toward Collaborative Intelligent WorldabstractCollaborative Intelligence is a new paradigm that splits a deep neural network (DNN) into an edge and cloud for deploying a DNN-based image recognition application. In this paradigm, deep features, which are the outputs of the edge DNN, are compressed and transmitted to the cloud DNN. Because the deep features have a number of responses that are similar to each other, for efficient compression, previous methods spatially arrange and compress the deep features as an image to utilize the similarity as a spatial correlation. However, if the deep features are arranged in not only spatial but also temporal directions like those in a video, it may be possible to compress them more efficiently by increasing a temporal correlation. To explore this possibility, we propose a “spatio-temporal arrangement”. This method spatially arranges the deep features as images and temporally arranges them as a video with a novel ordering search algorithm. Our method effectively increases the spatial and temporal correlations hidden in the deep features and achieves high compression efficiency compared with the previous methods. Experimental results demonstrate the compression efficiency of our method is better than that of the previous methods (1.50% to 4.98% on BD-Rate evaluation in a lossy setting). Our analysis shows that our method effectively increases the correlation when the input is an image with rich edges and textures. Shoichiro Takeda, Motohiro Takagi, Ryuichi Tanida, Hideaki Kimata, Hayaru Shouno |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Action Quality Assessment With Ignoring Scene ContextabstractWe propose an action quality assessment (AQA) method that can specifically assess target action quality with ignoring scene context, which is a feature unrelated to the target action. Existing AQA methods have tried to extract spatiotemporal features related to the target action by applying 3D convolution to the video. However, since their models are not explicitly designed to extract the features of the target action, they mis-extract scene context and thus cannot assess the target action quality correctly. To overcome this problem, we impose two losses to an existing AQA model: scene adversarial loss and our newly proposed human-masked regression loss. The scene adversarial loss encourages the model to ignore scene context by adversarial training. Our human-masked regression loss does so by making the correlation between score outputs by an AQA model and human referees undefinable when the target action is not visible. These two losses lead the model to specifically assess the target action quality with ignoring scene context. We evaluated our method on a diving dataset commonly used for AQA and found that it outperformed current state-of-the-art methods. This result shows that our method is effective in ignoring scene context while assessing the target action quality. Takasuke Nagai, Shoichiro Takeda, Masaaki Matsumura, Shinya Shimizu, Susumu Yamamoto |
ICIP | 2 |
| 2021 | Knowledge Transferred Fine-Tuning for Anti-Aliased Convolutional Neural Network in Data-Limited SituationabstractAnti-aliased convolutional neural networks (CNNs) introduce blur filters to intermediate representations in CNNs to achieve high accuracy. A promising way to build a new antialiased CNN is to fine-tune a pre-trained CNN, which can easily be found online, with blur filters. However, blur filters drastically degrade the pre-trained representation, so the fine-tuning needs to rebuild the representation by using massive training data. Therefore, if the training data is limited, the fine-tuning cannot work well because it induces overfitting to the limited training data. To tackle this problem, this paper proposes “knowledge transferred fine-tuning”. On the basis of the idea of knowledge transfer, our method transfers the knowledge from intermediate representations in the pre-trained CNN to the anti-aliased CNN while fine-tuning. We transfer only essential knowledge using a pixel-level loss that transfers detailed knowledge and a global-level loss that transfers coarse knowledge. Experimental results demonstrate that our method significantly outperforms the simple fine-tuning method. Shoichiro Takeda, Ryuichi Tanida, Hideaki Kimata, Hayaru Shouno |
ICIP | 2 |
| 2021 | Image Enhanced Rotation Prediction for Self-Supervised LearningabstractThe rotation prediction (Rotation) is a simple pretext-task for self-supervised learning (SSL), where models learn useful representations for target vision tasks by solving pretext-tasks. Although Rotation captures information of object shapes, it hardly captures information of textures. To tackle this problem, we introduce a novel pretext-task called image enhanced rotation prediction (IE-Rot) for SSL. IE-Rot simultaneously solves Rotation and another pretext-task based on image enhancement (e.g., sharpening and solarizing) while maintaining simplicity. Through the simultaneous prediction of rotation and image enhancement, models learn representations to capture the information of not only object shapes but also textures. Our experimental results show that IE-Rot models outperform Rotation on various standard benchmarks including ImageNet classification, PASCAL-VOC detection, and COCO detection/segmentation. Shin'ya Yamaguchi, Sekitoshi Kanai, Tetsuya Shioda, Shoichiro Takeda |
ICIP | 4 |
| 2020 | Deep Feature Compression With Spatio-Temporal Arranging for Collaborative IntelligenceabstractCollaborative Intelligence is a new paradigm that splits a deep neural network (DNN) into an edge DNN and a cloud DNN. In this paradigm, deep features, which are the outputs of the edge DNN, are compressed and transmitted to the cloud DNN. Since the deep features have a few responses that are similar to each other, previous studies have proposed compressing them as an image with a spatial arrangement to utilize spatial correlation between the deep features. However, this method may not sufficiently consider the similarity because only the spatial correlation is utilized. In this work, we propose a “spatio-temporal arranging” that considers the similarity of deep features in both the spatial and temporal directions. This method arranges the deep features spatiotemporally and compresses them as a video to utilize the spatio-temporal correlation. We perform spatial arranging as images to increase the spatial correlation and temporal arranging as a video with a novel ordering search to increase the temporal correlation. Experimental results demonstrate that our method performs better than previous methods. Motohiro Takagi, Shoichiro Takeda, Ryuichi Tanida, Hideaki Kimata |
ICIP | 3 |
| 2019 | Video Magnification in the Wild Using Fractional Anisotropy in Temporal DistributionabstractVideo magnification methods can magnify and reveal subtle changes invisible to the naked eye. However, in such subtle changes, meaningful ones caused by physical and natural phenomena are mixed with non-meaningful ones caused by photographic noise. Therefore, current methods often produce noisy and misleading magnification outputs due to the non-meaningful subtle changes. For detecting only meaningful subtle changes, several methods have been proposed but require human manipulations, additional resources, or input video scene limitations. In this paper, we present a novel method using fractional anisotropy (FA) to detect only meaningful subtle changes without the aforementioned requirements. FA has been used in neuroscience to evaluate anisotropic diffusion of water molecules in the body. On the basis of our observation that temporal distribution of meaningful subtle changes more clearly indicates anisotropic diffusion than that of non-meaningful ones, we used FA to design a fractional anisotropic filter that passes only meaningful subtle changes. Using the filter enables our method to obtain better and more impressive magnification results than those obtained with state-of-the-art methods. Shoichiro Takeda, Yasunori Akagi, Kazuki Okami, Megumi Isogai, Hideaki Kimata |
CVPR | 1 |
| 2018 | Jerk-Aware Video Acceleration MagnificationabstractVideo magnification reveals subtle changes invisible to the naked eye, but such tiny yet meaningful changes are often hidden under large motions: small deformation of the muscles in doing sports, or tiny vibrations of strings in ukulele playing. For magnifying subtle changes under large motions, video acceleration magnification method has recently been proposed. This method magnifies subtle acceleration changes and ignores slow large motions. However, quick large motions severely distort this method. In this paper, we present a novel use of jerk to make the acceleration method robust to quick large motions. Jerk has been used to assess smoothness of time series data in the neuroscience and mechanical engineering fields. On the basis of our observation that subtle changes are smoother than quick large motions at temporal scale, we used jerk-based smoothness to design a jerk-aware filter that passes subtle changes only under quick large motions. By applying our filter to the acceleration method, we obtain impressive magnification results better than those obtained with state-of-the-art. Shoichiro Takeda, Kazuki Okami, Dan Mikami, Megumi Isogai, Hideaki Kimata |
CVPR | 1 |