EDBT 2026 Demo / reviewers in the wild / expert
Masanori Suganuma
dblp:179/9075
· DBLP profile ↗
30ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-1469-9663ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RP-SLAM: Real-Time Photorealistic SLAM With Efficient 3D Gaussian Splattingabstract3D Gaussian Splatting has emerged as a promising technique for high-quality 3D rendering, leading to increasing interest in integrating 3DGS into realism SLAM systems. However, existing methods face challenges such as Gaussian primitives redundancy, forgetting problem during continuous optimization, and difficulty in initializing primitives in monocular case due to lack of depth information. In order to achieve efficient and photorealistic mapping, we propose RP-SLAM, a 3D Gaussian splatting-based vision SLAM method for monocular and RGB-D cameras. RP-SLAM decouples camera poses estimation from Gaussian primitives optimization and consists of three key components. Firstly, we propose an efficient incremental mapping approach to achieve a compact and accurate representation of the scene through adaptive sampling and Gaussian primitives filtering. Secondly, a dynamic window optimization method is proposed to mitigate the forgetting problem and improve map consistency. Finally, for the monocular case, a monocular keyframe initialization method based on sparse point cloud is proposed to improve the initialization accuracy of Gaussian primitives, which provides a geometric basis for subsequent optimization. The results of numerous experiments demonstrate that RP-SLAM achieves state-of-the-art map rendering accuracy while ensuring real-time performance and model compactness. Lizhi Bai, Chunqi Tian, Jun Yang 0056, Masanori Suganuma, Takayuki Okatani |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Self-Supervised Learning of Intertwined Content and Positional Features for Object DetectionabstractWe present a novel self-supervised feature learning method using Vision Transformers (ViT) as the backbone, specifically designed for object detection and instance segmentation. Our approach addresses the challenge of extracting features that capture both class and positional information, which are crucial for these tasks. The method introduces two key components: (1) a positional encoding tied to the cropping process in contrastive learning, which utilizes a novel vector field representation for positional embeddings; and (2) masking and prediction, similar to conventional Masked Image Modeling (MIM), applied in parallel to both content and positional embeddings of image patches. These components enable the effective learning of intertwined content and positional features. We evaluate our method against state-of-the-art approaches, pre-training on ImageNet-1K and fine-tuning on downstream tasks. Our method outperforms the state-of-the-art SSL methods on the COCO object detection benchmark, achieving significant improvements with fewer pre-training epochs. These results suggest that better integration of positional information into self-supervised learning can improve performance on the dense prediction tasks. Kang-Jun Liu, Masanori Suganuma, Takayuki Okatani |
ICML | 2 |
| 2025 | MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image RetrievalabstractResult diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method. Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Masanori Suganuma, Takayuki Okatani |
IJCAI | 4 |
| 2025 | Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel MethodabstractThis paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a generated image. Existing methods often struggle with accuracy in this task. We propose a novel hybrid approach that combines direct inversion via gradient descent for the first step, followed by a fixed-point method for subsequent steps. Empirical evaluations across three datasets demonstrate that our method significantly improves the prediction of initial latent variables while achieving superior reconstruction accuracy. Additionally, we introduce a new evaluation, called the self-interpolation test, which assesses the quality of images generated from interpolated points between the true and predicted latent maps, offering deeper insights into performance. Our results reveal that while existing methods perform reasonably well in reconstruction, they consistently fail to accurately predict the initial latent variables, resulting in poor performance on the self-interpolation test. In contrast, our method outperforms all others across all metrics, providing valuable insights into diffusion models and enhancing their applications in image generation and editing. Masanori Suganuma, Takayuki Okatani |
WACV | 2 |
| 2025 | RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolutionabstractSmartphones with multi-camera systems, featuring cameras with varying field-of-views (FoVs), are increasingly common. This variation in FoVs results in content differences across videos, paving the way for an innovative approach to video super-resolution (VSR). This method enhances the VSR performance of lower resolution (LR) videos by leveraging higher resolution reference (Ref) videos. Previous works [14, 15 ], which operate on this principle, generally expand on traditional VSR models by combining LR and Ref inputs over time into a unified stream. However, we can expect that better results are obtained by independently aggre-gating these Ref image sequences temporally. Therefore, we introduce an improved method, RefVSR++, which performs the parallel aggregation of LR and Ref images in the temporal direction, aiming to optimize the use of the available data. RefVSR++ also incorporates improved mechanisms for aligning image features over time, crucial for effective VSR. Our experiments demonstrate that RefVSR++ outper-forms previous works by over 1dB in PSNR, setting a new benchmark in the field. Han Zou, Masanori Suganuma, Takayuki Okatani |
WACV | 2 |
| 2025 | Rethinking Open-Set Object Detection: Issues, A New Formulation, and TaxonomyabstractAbstract Open-set object detection (OSOD), a task involving the detection of unknown objects while accurately detecting known objects, has recently gained attention. However, we identify a fundamental issue with the problem formulation employed in current OSOD studies. Inherent to object detection is knowing “what to detect,” which contradicts the idea of identifying “unknown” objects. This sets OSOD apart from open-set recognition (OSR). This contradiction complicates a proper evaluation of methods’ performance, a fact that previous studies have overlooked. Next, we propose a novel formulation wherein detectors are required to detect both known and unknown classes within specified super-classes of object classes. This new formulation is free from the aforementioned issues and has practical applications. Finally, we design benchmark tests utilizing existing datasets and report the experimental evaluation of existing OSOD methods. The results show that existing methods fail to accurately detect unknown objects due to misclassification of known and unknown classes rather than incorrect bounding box prediction. As a byproduct, we introduce a taxonomy of OSOD, resolving confusion prevalent in the literature. We anticipate that our study will encourage the research community to reconsider OSOD and facilitate progress in the right direction. Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani |
Int. J. Comput. Vis. | 2 |
| 2024 | SBCFormer: Lightweight Network Capable of Full-size ImageNet Classification at 1 FPS on Single Board ComputersabstractComputer vision has become increasingly prevalent in solving real-world problems across diverse domains, including smart agriculture, fishery, and livestock management. These applications may not require processing many image frames per second, leading practitioners to use single board computers (SBCs). Although many lightweight networks have been developed for "mobile/edge" devices, they primarily target smartphones with more powerful processors and not SBCs with the low-end CPUs. This paper introduces a CNN-ViT hybrid network called SBCFormer, which achieves high accuracy and fast computation on such low-end CPUs. The hardware constraints of these CPUs make the Transformer’s attention mechanism preferable to convolution. However, using attention on low-end CPUs presents a challenge: high-resolution internal feature maps demand excessive computational resources, but reducing their resolution results in the loss of local image details. SBCFormer introduces an architectural design to address this issue. As a result, SBCFormer achieves the highest trade-off between accuracy and speed on a Raspberry Pi 4 Model B with an ARM-Cortex A72 CPU. For the first time, it achieves an ImageNet-1K top-1 accuracy of around 80% at a speed of 1.0 frame/sec on the SBC. Code is available at https://github.com/xyongLu/SBCFormer. Xiangyong Lu, Masanori Suganuma, Takayuki Okatani |
WACV | 2 |
| 2024 | Contextual Affinity Distillation for Image Anomaly DetectionabstractPrevious studies on unsupervised industrial anomaly detection mainly focus on ‘structural’ types of anomalies such as cracks and color contamination by matching or learning local feature representations. While achieving significantly high detection performance on this kind of anomaly, they are faced with ‘logical’ types of anomalies that violate the long-range dependencies such as a normal object placed in the wrong position. Noting the reverse distillation approaches that are under the encoder-decoder paradigm could learn from the high abstract level knowledge, we propose to use two students (local and global) to better mimic the teacher’s local and global behavior in reverse distillation. The local student, which is used in previous studies mainly focuses on accurate local feature learning while the global student pays attention to learning global correlations. To further encourage the global student’s learning to capture long-range dependencies, we design the global context condensing block (GCCB) and propose a contextual affinity loss for the student training and anomaly scoring. Experimental results show that the proposed method sets a new state-of-the-art performance on the MVTec LOCO AD dataset without using complex training techniques. Masanori Suganuma, Takayuki Okatani |
WACV | 2 |
| 2024 | Improved high dynamic range imaging using multi-scale feature flows balanced between task-orientedness and accuracyabstractDeep learning has made it possible to accurately generate high dynamic range (HDR) images from multiple images taken at different exposure settings, largely owing to advancements in neural network design. However, generating images without artifacts remains difficult, especially in scenes with moving objects. In such cases, issues like color distortion, geometric misalignment, or ghosting can appear. Current state-of-the-art network designs address this by estimating the optical flow between input images to align them better. The parameters for the flow estimation are learned through the primary goal, producing high-quality HDR images. However, we find that this ”task-oriented flow” approach has its drawbacks, especially in minimizing artifacts. To address this, we introduce a new network design and training method that improve the accuracy of flow estimation. This aims to strike a balance between task-oriented flow and accurate flow. Additionally, the network utilizes multi-scale features extracted from the input images for both flow estimation and HDR image reconstruction. Our experiments demonstrate that these two innovations result in HDR images with fewer artifacts and enhanced quality. Masanori Suganuma, Takayuki Okatani |
Comput. Vis. Image Underst. | 2 |
| 2024 | Symmetry-aware Neural Architecture for Embodied Visual NavigationabstractAbstract The existing methods for addressing visual navigation employ deep reinforcement learning as the standard tool for the task. However, they tend to be vulnerable to statistical shifts between the training and test data, resulting in poor generalization over novel environments that are out-of-distribution from the training data. In this study, we attempt to improve the generalization ability by utilizing the inductive biases available for the task. Employing the active neural SLAM that learns policies with the advantage actor-critic method as the base framework, we first point out that the mappings represented by the actor and the critic should satisfy specific symmetries. We then propose a network design for the actor and the critic to inherently attain these symmetries. Specifically, we use G-convolution instead of the standard convolution and insert the semi-global polar pooling layer, which we newly design in this study, in the last section of the critic network. Our method can be integrated into existing methods that utilize intermediate goals and 2D occupancy maps. Experimental results show that our method improves generalization ability by a good margin over visual exploration and object goal navigation, which are two main embodied visual navigation tasks. Shuang Liu 0002, Masanori Suganuma, Takayuki Okatani |
Int. J. Comput. Vis. | 2 |
| 2024 | That's BAD: blind anomaly detection by implicit local feature clusteringabstractAbstract Recent studies on visual anomaly detection (AD) of industrial objects/textures have achieved quite good performance. They consider an unsupervised setting, specifically the one-class setting, in which we assume the availability of a set of normal (i.e., anomaly-free) images for training. In this paper, we consider a more challenging scenario of unsupervised AD, in which we detect anomalies in a given set of images that might contain both normal and anomalous samples. The setting does not assume the availability of known normal data and thus is completely free from human annotation, which differs from the standard AD considered in recent studies. For clarity, we call the setting blind anomaly detection (BAD). We show that BAD can be converted into a local outlier detection problem and propose a novel method named PatchCluster that can accurately detect image- and pixel-level anomalies. Experimental results show that PatchCluster shows a promising performance without the knowledge of normal data, even comparable to the SOTA methods applied in the one-class setting needing it. Masanori Suganuma, Takayuki Okatani |
Mach. Vis. Appl. | 2 |
| 2024 | Rethinking unsupervised domain adaptation for semantic segmentationabstractUnsupervised domain adaptation (UDA) adapts a model trained on one domain (called source) to a novel domain (called target) using only unlabeled data. Due to its high annotation cost, researchers have developed many UDA methods for semantic segmentation, which assume no labeled sample is available in the target domain. We question the practicality of this assumption for two reasons. First, after training a model with a UDA method, we must somehow verify the model before deployment. Second, UDA methods have at least a few hyper-parameters that need to be determined. The surest solution to these is to evaluate the model using validation data, i.e., a certain amount of labeled target-domain samples. This question about the basic assumption of UDA leads us to rethink UDA from a data-centric point of view. Specifically, we assume we have access to a minimum level of labeled data. Then, we ask how much is necessary to find good hyper-parameters of existing UDA methods. We then consider what if we use the same data for supervised training of the same model, e.g., finetuning. We conducted experiments to answer these questions with popular scenarios, {GTA5, SYNTHIA} → Cityscapes. We found that i) choosing good hyper-parameters needs only a few labeled images for some UDA methods whereas a lot more for others; and ii) simple finetuning works surprisingly well; it outperforms many UDA methods if only several dozens of labeled images are available. • We rethink the UDA for segmentation from a data-centric perspective. • Our starting point is that any ML system requires annotated data for validation. • We investigate how much data is necessary to select parameters of existing methods. • We consider what if we use the same data for supervised training of the same model. Masanori Suganuma, Takayuki Okatani |
Pattern Recognit. Lett. | 2 |
| 2023 | Accurate Single-Image Defocus Deblurring Based on Improved Integration with Defocus Map EstimationabstractThis paper considers the problem of single-image defocus deblurring, which involves removing blur in an input image caused by defocusing. Previous studies have employed two main approaches, the first being a two-step approach involving estimating the defocus map from the input image and then computing the blur kernel from it, followed by non-blind deconvolution to obtain the estimate of the clean image. The second approach is a direct method where the clean image is estimated directly from the blurry input image. The paper proposes an intermediate approach that explicitly estimates the defocus map of the scene but does not explicitly compute the kernel or its inverse. Instead, it attempts to learn a direct mapping from the blurry input image to the clean image by utilizing the estimated defocus map to condition the mapping. Experimental results show that the proposed method can yield higher quality outputs than the state-of-the-art methods. Masanori Suganuma, Takayuki Okatani |
ICIP | 2 |
| 2023 | Network Pruning and Fine-tuning for Few-shot Industrial Image Anomaly DetectionabstractThis paper focuses on industrial image anomaly detection and localization under few-shot settings. Since acquiring sufficient anomalous data is difficult, unsupervised learning that uses only normal data is commonly used, but even obtaining enough anomaly-free training samples can be challenging. Moreover, applying data augmentations, which is a common strategy for few-shot learning to alleviate the lack of data, is limited to use for some industrial product images. To address the above issues, we propose a network pruning and fine-tuning (PF) framework that leverages the knowledge of a deep pre-trained model. Our approach distills the knowledge of normal samples into a pruned student network, followed by fine-tuning to restore its representation ability for normal data. During inference, discrepancies between features extracted by the teacher and student are used to determine the anomaly score. The proposed method could better utilize the strong representation ability of deep models and benefit the student training with limited data by network pruning. Our framework achieves state-of-the-art performance on the MVTec AD benchmark and is not limited to specific network pruning methods. Masanori Suganuma, Takayuki Okatani |
INDIN | 2 |
| 2023 | Unsupervised domain adaptation for semantic segmentation via cross-region alignmentabstractSemantic segmentation requires a lot of training data, which necessitates costly annotation. There have been many studies on unsupervised domain adaptation (UDA) from one domain to another, e.g., from computer graphics to real images. However, there is still a gap in accuracy between UDA and supervised training on native domain data. It is arguably attributable to the class-level misalignment between the source and target domain data. To cope with this, we propose a method that applies adversarial training to align two feature distributions in the target domain. It uses a self-training framework to split the image into two regions (i.e., trusted and untrusted), which form two distributions to align in the feature space. We term this approach cross-region adaptation (CRA) to distinguish it from the previous methods of aligning different domain distributions, which we call cross-domain adaptation (CDA). CRA can be applied after any CDA method. Experimental results show that this always improves the accuracy of the combined CDA method. Xing Liu 0010, Masanori Suganuma, Takayuki Okatani |
Comput. Vis. Image Underst. | 3 |
| 2022 | GRIT: Faster and Better Image Captioning Transformer Using Dual Visual Features
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
ECCV (36) | 2 |
| 2022 | Bridging the Gap from Asymmetry Tricks to Decorrelation Principles in Non-contrastive Self-supervised LearningabstractRecent non-contrastive methods for self-supervised representation learning show promising performance. While they are attractive since they do not need negative samples, it necessitates some mechanism to avoid collapsing into a trivial solution. Currently, there are two approaches to collapse prevention. One uses an asymmetric architecture on a joint embedding of input, e.g., BYOL and SimSiam, and the other imposes decorrelation criteria on the same joint embedding, e.g., Barlow-Twins and VICReg. The latter methods have theoretical support from information theory as to why they can learn good representation. However, it is not fully understood why the former performs equally well. In this paper, focusing on BYOL/SimSiam, which uses the stop-gradient and a predictor as asymmetric tricks, we present a novel interpretation of these tricks; they implicitly impose a constraint that encourages feature decorrelation similar to Barlow-Twins/VICReg. We then present a novel non-contrastive method, which replaces the stop-gradient in BYOL/SimSiam with the derived constraint; the method empirically shows comparable performance to the above SOTA methods in the standard benchmark test using ImageNet. This result builds a bridge from BYOL/SimSiam to the decorrelation-based methods, contributing to demystifying their secrets. Kang-Jun Liu, Masanori Suganuma, Takayuki Okatani |
NeurIPS | 2 |
| 2021 | Matching in the Dark: A Dataset for Matching Image Pairs of Low-light ScenesabstractThis paper considers matching images of low-light scenes, aiming to widen the frontier of SfM and visual SLAM applications. Recent image sensors can record the brightness of scenes with more than eight-bit precision, available in their RAW-format image. We are interested in making full use of such high-precision information to match extremely low-light scene images that conventional methods cannot handle. For extreme low-light scenes, even if some of their brightness information exists in the RAW format images’ low bits, the standard raw image processing on cameras fails to utilize them properly. As was recently shown by Chen et al. [14], CNNs can learn to produce images with a natural appearance from such RAW-format images. To consider if and how well we can utilize such information stored in RAW-format images for image matching, we have created a new dataset named MID (matching in the dark). Using it, we experimentally evaluated combinations of eight image-enhancing methods and eleven image matching methods consisting of classical/neural local descriptors and classical/neural initial point-matching methods. The results show the advantage of using the RAW-format images and the strengths and weaknesses of the above component methods. They also imply there is room for further research. Wenzheng Song, Masanori Suganuma, Xing Liu 0010, Noriyuki Shimobayashi, Daisuke Maruta, Takayuki Okatani |
ICCV | 2 |
| 2021 | Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following TasksabstractThere is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives. Recent studies have tackled the problem using ALFRED, a well-designed dataset for the task, but achieved only very low accuracy. This paper proposes a new method, which outperforms the previous methods by a large margin. It is based on a combination of several new ideas. One is a two-stage interpretation of the provided instructions. The method first selects and interprets an instruction without using visual information, yielding a tentative action sequence prediction. It then integrates the prediction with the visual information etc., yielding the final prediction of an action and an object. As the object's class to interact is identified in the first stage, it can accurately select the correct object from the input image. Moreover, our method considers multiple egocentric views of the environment and extracts essential information by applying hierarchical attention conditioned on the current instruction. This contributes to the accurate prediction of actions for navigation. A preliminary version of the method won the ALFRED Challenge 2020. The current version achieves the unseen environment's success rate of 4.45% with a single view, which is further improved to 8.37% with multiple views. Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
IJCAI | 2 |
| 2020 | Hyperparameter-Free Out-of-Distribution Detection Using Cosine Similarity
Engkarat Techapanurak, Masanori Suganuma, Takayuki Okatani |
ACCV (4) | 2 |
| 2020 | Efficient Attention Mechanism for Visual Dialog that Can Handle All the Interactions Between Multiple Inputs
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
ECCV (24) | 2 |
| 2020 | Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object DetectorsabstractThe employment of convolutional neural networks has led to significant performance improvement on the task of object detection. However, when applying existing detectors to continuous frames in a video, we often encounter momentary miss-detection of objects, that is, objects are undetected exceptionally at a few frames, although they are correctly detected at all other frames. In this paper, we analyze the mechanism of how such miss-detection occurs. For the most popular class of detectors that are based on anchor boxes, we show the followings: i) besides apparent causes such as motion blur, occlusions, background clutters, etc., the majority of remaining miss-detection can be explained by an improper behavior of the detectors at boundaries of the anchor boxes; and ii) this can be rectified by improving the way of choosing positive samples from candidate anchor boxes when training the detectors. Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani |
WACV | 2 |
| 2020 | Evolution of Deep Convolutional Neural Networks Using Cartesian Genetic ProgrammingabstractThe convolutional neural network (CNN), one of the deep learning models, has demonstrated outstanding performance in a variety of computer vision tasks. However, as the network architectures become deeper and more complex, designing CNN architectures requires more expert knowledge and trial and error. In this article, we attempt to automatically construct high-performing CNN architectures for a given task. Our method uses Cartesian genetic programming (CGP) to encode the CNN architectures, adopting highly functional modules such as a convolutional block and tensor concatenation, as the node functions in CGP. The CNN structure and connectivity, represented by the CGP, are optimized to maximize accuracy using the evolutionary algorithm. We also introduce simple techniques to accelerate the architecture search: rich initialization and early network training termination. We evaluated our method on the CIFAR-10 and CIFAR-100 datasets, achieving competitive performance with state-of-the-art models. Remarkably, our method can find competitive architectures with a reasonable computational cost compared to other automatic design methods that require considerably more computational time and machine resources. Masanori Suganuma, Masayuki Kobayashi, Shinichi Shirakawa, Tomoharu Nagao |
Evol. Comput. | 1 |
| 2019 | Dual Residual Networks Leveraging the Potential of Paired Operations for Image RestorationabstractIn this paper, we study design of deep neural networks for tasks of image restoration. We propose a novel style of residual connections dubbed "dual residual connection", which exploits the potential of paired operations, e.g., up- and down-sampling or convolution with large- and small-size kernels. We design a modular block implementing this connection style; it is equipped with two containers to which arbitrary paired operations are inserted. Adopting the "unraveled" view of the residual networks proposed by Veit et al., we point out that a stack of the proposed modular blocks allows the first operation in a block interact with the second operation in any subsequent blocks. Specifying the two operations in each of the stacked blocks, we build a complete network for each individual task of image restoration. We experimentally evaluate the proposed approach on five image restoration tasks using nine datasets. The results show that the proposed networks with properly chosen paired operations outperform previous methods on almost all of the tasks and datasets. Xing Liu 0010, Masanori Suganuma, Zhun Sun, Takayuki Okatani |
CVPR | 2 |
| 2019 | Attention-Based Adaptive Selection of Operations for Image Restoration in the Presence of Unknown Combined DistortionsabstractMany studies have been conducted so far on image restoration, the problem of restoring a clean image from its distorted version. There are many different types of distortion affecting image quality. Previous studies have focused on single types of distortion, proposing methods for removing them. However, image quality degrades due to multiple factors in the real world. Thus, depending on applications, e.g., vision for autonomous cars or surveillance cameras, we need to be able to deal with multiple combined distortions with unknown mixture ratios. For this purpose, we propose a simple yet effective layer architecture of neural networks. It performs multiple operations in parallel, which are weighted by an attention mechanism to enable selection of proper operations depending on the input. The layer can be stacked to form a deep network, which is differentiable and thus can be trained in an end-to-end fashion by gradient descent. The experimental results show that the proposed method works better than previous methods by a good margin on tasks of restoring images with multiple combined distortions. Masanori Suganuma, Xing Liu 0010, Takayuki Okatani |
CVPR | 1 |
| 2018 | Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary SearchabstractResearchers have applied deep neural networks to image restoration tasks, in which they proposed various network architectures, loss functions, and training methods. In particular, adversarial training, which is employed in recent studies, seems to be a key ingredient to success. In this paper, we show that simple convolutional autoencoders (CAEs) built upon only standard network components, i.e., convolutional layers and skip connections, can outperform the state-of-the-art methods which employ adversarial training and sophisticated loss functions. The secret is to search for good architectures using an evolutionary algorithm. All we did was to train the optimized CAEs by minimizing the l2 loss between reconstructed images and their ground truths using the ADAM optimizer. Our experimental results show that this approach achieves 27.8 dB peak signal to noise ratio (PSNR) on the CelebA dataset and 33.3 dB on the SVHN dataset, compared to 22.8 dB and 19.0 dB provided by the former state-of-the-art methods, respectively. Masanori Suganuma, Mete Ozay, Takayuki Okatani |
ICML | 1 |
| 2018 | A Genetic Programming Approach to Designing Convolutional Neural Network ArchitecturesabstractWe propose a method for designing convolutional neural network (CNN) architectures based on Cartesian genetic programming (CGP). In the proposed method, the architectures of CNNs are represented by directed acyclic graphs, in which each node represents highly-functional modules such as convolutional blocks and tensor operations, and each edge represents the connectivity of layers. The architecture is optimized to maximize the classification accuracy for a validation dataset by an evolutionary algorithm. We show that the proposed method can find competitive CNN architectures compared with state-of-the-art methods on the image classification task using CIFAR-10 and CIFAR-100 datasets. Masanori Suganuma, Shinichi Shirakawa, Tomoharu Nagao |
IJCAI | 1 |
| 2017 | A genetic programming approach to designing convolutional neural network architecturesabstractThe convolutional neural network (CNN), which is one of the deep learning models, has seen much success in a variety of computer vision tasks. However, designing CNN architectures still requires expert knowledge and a lot of trial and error. In this paper, we attempt to automatically construct CNN architectures for an image classification task based on Cartesian genetic programming (CGP). In our method, we adopt highly functional modules, such as convolutional blocks and tensor concatenation, as the node functions in CGP. The CNN structure and connectivity represented by the CGP encoding method are optimized to maximize the validation accuracy. To evaluate the proposed method, we constructed a CNN architecture for the image classification task with the CIFAR-10 dataset. The experimental result shows that the proposed method can be used to automatically find the competitive CNN architecture compared with state-of-the-art models. Masanori Suganuma, Shinichi Shirakawa, Tomoharu Nagao |
GECCO | 1 |
| 2017 | Acquiring grasp strategies for a multifingered robot hand using evolutionary algorithmsabstractIn recent years, significant research has been conducted on grasp planning for multifingered robot hands. These studies have focused on determining how to obtain suitable grasps from among an infinite number of candidate grasps. This domain's goal is a successful application to unknown environments through the adoption of the extracted grasps. Under difficult conditions, such as grasping a target object that is adjacent to other objects, manipulating robot hands by indicating grasping points has been insufficient. Instead, grasp strategies that construct movements using each finger's joint servo controls and robot hand movements should be used. In addition, it is necessary to automatically acquire various grasp strategies to apply to unknown environments. In this paper, we propose a method that automatically obtains grasp strategies using a real-coded genetic algorithm (RCGA), which is an evolutionary algorithm. This method derives grasp strategies by optimizing combinations and structures that consist of simple finger joint servo controls and robot hand movements. By applying our method to several objects on a simulator, we collected various grasp strategies capable of handling difficult conditions. Chiaki Hirayama, Toshiya Watanabe, Shinji Kawabata, Masanori Suganuma, Tomoharu Nagao |
SMC | 4 |
| 2016 | Hierarchical feature construction for image classification using Genetic ProgrammingabstractIn this paper, we design a hierarchical feature construction method for image classification. Our method has two feature construction stages: (1) feature construction by a combination of primitive image processing filters, and (2) feature construction by evolved filters. We verify the image classification performance of the proposed method on the MIT urban and nature scene dataset. The experimental results show that the two-stage feature construction improves the classification accuracy compared to single stage feature construction. In addition, the proposed method outperforms several existing feature construction methods. Masanori Suganuma, Daiki Tsuchiya, Shinichi Shirakawa, Tomoharu Nagao |
SMC | 1 |